Nauto+Nexar

Test it yourself

Four things you can check today without speaking to anyone here. The model is open, the scorecard is published, and the method is written down.

[ FOUR THINGS YOU CAN CHECK ]

How to put Empiric Earth to the test

[ Disclaimer ]

Before you benchmark the open model, read this first.

It is not the model we deploy and it scores lower, on a fraction of the training data. It is published so the approach can be inspected and run locally, not as a stand-in for what ships. The numbers for both are below.

[ SCOPED TO ALERT TESTING ]

A test of alerting, not a benchmark of a perception model.

Controlled testing by the Virginia Tech Transportation Institute. The published Phase 1 report gives 100% alert accuracy across the tested distraction behaviours and an average time to alert of 3.8 seconds.

Disclosure: Empiric Earth funded the December 2024 evaluation. VTTI does not endorse any products. We state this wherever the figures appear, not only where it is convenient.

[ PUBLISHED MODEL BENCHMARKS ]

Our own test, with the definitions attached.

BADAS publishes its results across its own long-tail benchmark categories, along with the definitions those categories use, and external comparisons are described separately. It is trained on real edge-device footage with zero synthetic data.

This is an internal benchmark, published openly. It is not independent testing, and we do not describe it as such.

DAD · BADAS, deployed

AUC 0.992 · AP 0.920

DAD · BADAS-open

AUC 0.870 · AP 0.660

DADA-2000 · BADAS, deployed

AUC 0.991 · AP 0.996

DADA-2000 · BADAS-open

AUC 0.770 · AP 0.870

Open

The model is downloadable on Hugging Face. You can run it against your own held-out data and disagree with us in public.

Defined

Every benchmark category is published with the definition it uses, so the boundary is inspectable rather than convenient.

Zero synthetic

Trained on real edge-device footage. No generated training data, which is the claim the rest of the argument rests on.

[ Publicly Available Assets ]

Everything else is already published.

Open model weights, Apache 2.0

huggingface.co/nexar-ai/BADAS-Open

The dataset

1,500 clips of about 40 seconds of real driving, labelled with what happened, the lighting, the weather and the road type.

The dataset paper

Moura, Zhu and Zvitia, March 2025. arxiv.org/abs/2503.03848

The Kaggle competition

Where other people’s models were scored on it under a fixed protocol.

Two peer-reviewed studies

On vulnerable road users, with Waymo.

What the model was trained on

178,500 labelled videos, roughly 2M windowed clips, all real driving from edge devices. No synthetic data.

What it was tested on

The Kaggle competition set of 1,344 clips, a ten-group long-tail benchmark, and the public DAD and DADA-2000 sets, which other research groups built.

[ INDEPENDENT TESTING · VTTI ]

The in-cab alerting system, tested by VTTI.

100%

alert accuracy

3.8 s

average time to alert

[ NAMED CUSTOMER OUTCOMES ]

Scoped to the deployment.

Each figure belongs to one customer, one fleet profile and one measured period. None of them is a universal claim about what the product does, and we will not present them as one.

Customer results reflect each fleet’s deployment, baseline, operating conditions and measurement period. The full published story carries that context. Every figure on this page is subject to a content validity check before publication.

[ NOT YET PUBLISHED ]

What we are not claiming.

A page like this is only credible if it also says what is missing. Four things we have been asked about and are not asserting.

Φ

A universal collision-reduction figure

Published results run from 62% to 100% across deployments, on measures that are not comparable with one another — at-fault collisions, total collisions, most-serious collisions and distraction events are four different things. A single blended percentage would require a method across unlike measures, so we quote deployments instead.

Δ

An edge-case count as a headline figure

The 60M+ edge-case total is defined in The Delta Numbers, and the definition is what defends it. Because the boundary is a judgement, we publish it with its definition and never lead with it.

Σ

A production OEM outcome

Automotive relationships are live. A measured result from a factory-fit deployment is not published yet, and will appear here with its scope when it is.

Ψ

Benchmark leadership across the board

Our model leads its published long-tail categories. Independent VTTI testing measured alert performance. Those are two claims and they stay two claims.

[ SCALE ]

Both companies began measuring in 2015 and neither stopped.

Six figures, each describing a different population. They are not interchangeable and we do not add them together.

Miles of driving history

Over 10 billion

Real-world miles a month

More than 300 million, across 50+ countries

US roads reached

98%

Active sensors

350,000

Objects detected daily

80M+

Edge cases kept

More than 60 million.