
+
Test it yourself
Four things you can check today without speaking to anyone here. The model is open, the scorecard is published, and the method is written down.
[ FOUR THINGS YOU CAN CHECK ]
How to put Empiric Earth to the test
Φ
Run BADAS on your own video
Upload a clip and watch the collision probability move frame by frame, including where the model gets it wrong.
Try BADAS →
Δ
Download the open model
Open weights under Apache 2.0. Run it on your own hardware against your own held-out data.
Visit huggingface →
Σ
Query Risk Index
Road risk for somewhere you already know, weighted by the traffic actually passing through. A place you know well tells you quickly whether we are describing your roads correctly.
Query Risk Index →
Ψ
Have your own system scored by APEX
AV readiness. We benchmark an AV or ADAS system against a curated long-tail set it has not seen, and report where it stands.
Try APEX →

[ Disclaimer ]
Before you benchmark the open model, read this first.
It is not the model we deploy and it scores lower, on a fraction of the training data. It is published so the approach can be inspected and run locally, not as a stand-in for what ships. The numbers for both are below.
[ SCOPED TO ALERT TESTING ]
A test of alerting, not a benchmark of a perception model.
Controlled testing by the Virginia Tech Transportation Institute. The published Phase 1 report gives 100% alert accuracy across the tested distraction behaviours and an average time to alert of 3.8 seconds.
Disclosure: Empiric Earth funded the December 2024 evaluation. VTTI does not endorse any products. We state this wherever the figures appear, not only where it is convenient.
[ PUBLISHED MODEL BENCHMARKS ]
Our own test, with the definitions attached.
BADAS publishes its results across its own long-tail benchmark categories, along with the definitions those categories use, and external comparisons are described separately. It is trained on real edge-device footage with zero synthetic data.
This is an internal benchmark, published openly. It is not independent testing, and we do not describe it as such.
DAD · BADAS, deployed
AUC 0.992 · AP 0.920
DAD · BADAS-open
AUC 0.870 · AP 0.660
DADA-2000 · BADAS, deployed
AUC 0.991 · AP 0.996
DADA-2000 · BADAS-open
AUC 0.770 · AP 0.870
Open
The model is downloadable on Hugging Face. You can run it against your own held-out data and disagree with us in public.
Defined
Every benchmark category is published with the definition it uses, so the boundary is inspectable rather than convenient.
Zero synthetic
Trained on real edge-device footage. No generated training data, which is the claim the rest of the argument rests on.
[ Publicly Available Assets ]
Everything else is already published.
Open model weights, Apache 2.0
The dataset
1,500 clips of about 40 seconds of real driving, labelled with what happened, the lighting, the weather and the road type.
The dataset paper
Moura, Zhu and Zvitia, March 2025. arxiv.org/abs/2503.03848
The Kaggle competition
Where other people’s models were scored on it under a fixed protocol.
Two peer-reviewed studies
On vulnerable road users, with Waymo.
What the model was trained on
178,500 labelled videos, roughly 2M windowed clips, all real driving from edge devices. No synthetic data.
What it was tested on
The Kaggle competition set of 1,344 clips, a ten-group long-tail benchmark, and the public DAD and DADA-2000 sets, which other research groups built.
[ INDEPENDENT TESTING · VTTI ]
The in-cab alerting system, tested by VTTI.
100%
alert accuracy
3.8 s
average time to alert
[ NAMED CUSTOMER OUTCOMES ]
Scoped to the deployment.
Each figure belongs to one customer, one fleet profile and one measured period. None of them is a universal claim about what the product does, and we will not present them as one.
Customer results reflect each fleet’s deployment, baseline, operating conditions and measurement period. The full published story carries that context. Every figure on this page is subject to a content validity check before publication.
[ NOT YET PUBLISHED ]
What we are not claiming.
A page like this is only credible if it also says what is missing. Four things we have been asked about and are not asserting.
Φ
A universal collision-reduction figure
Published results run from 62% to 100% across deployments, on measures that are not comparable with one another — at-fault collisions, total collisions, most-serious collisions and distraction events are four different things. A single blended percentage would require a method across unlike measures, so we quote deployments instead.
Δ
An edge-case count as a headline figure
The 60M+ edge-case total is defined in The Delta Numbers, and the definition is what defends it. Because the boundary is a judgement, we publish it with its definition and never lead with it.
Σ
A production OEM outcome
Automotive relationships are live. A measured result from a factory-fit deployment is not published yet, and will appear here with its scope when it is.
Ψ
Benchmark leadership across the board
Our model leads its published long-tail categories. Independent VTTI testing measured alert performance. Those are two claims and they stay two claims.
[ SCALE ]
Both companies began measuring in 2015 and neither stopped.
Six figures, each describing a different population. They are not interchangeable and we do not add them together.
Miles of driving history
Over 10 billion
Real-world miles a month
More than 300 million, across 50+ countries
US roads reached
98%
Active sensors
350,000
Objects detected daily
80M+
Edge cases kept
More than 60 million.