CarKept

Methodology and data source

Data source

All statistics on this site are calculated from the Driver and Vehicle Standards Agency (DVSA) “Anonymised MOT tests and results” dataset, published on data.gov.uk under the Open Government Licence v3.0. This site contains public sector information licensed under the OGL.

This site’s dataset

Filters applied

Only test class 4vehicles are included. DVSA defines class 4 as “cars, passenger vehicles, motor caravans, private hire vehicles, motor tricycles, quadricycles and dual-purpose vehicles” up to eight passenger seats, plus goods vehicles up to 3,000kg and taxis/ambulances. This is the closest DVSA test class to “cars”, but the dataset has no separate body-type field, so a small number of light vans, motor caravans and taxis are included alongside cars within class 4.

Only normal (initial) testsare included (DVSA test type “NT”), excluding retests, partial retests and appeals. This matches DVSA’s own published effectiveness-report methodology, which uses “Normal (Initial) tests with outcomes of Pass, Fail or PRS” and omits all other test types.

Pass ratecounts both straightforward passes and “Pass with Rectification at Station” (PRS) — where a defect was fixed on the spot before the certificate was issued — as a pass, since the vehicle left with a valid MOT either way.

Failure reasonsare drawn from tests with a Fail or PRS outcome, using only test items marked as a Fail-type or PRS-type reason for rejection (excluding advisory notices and minor items that do not affect the test result). Each reason is mapped to its top-level DVSA test item category (e.g. “Brakes”, “Tyres”) using DVSA’s own item/RfR lookup tables.

Average mileageexcludes tests with a blank or zero mileage reading, which DVSA’s documentation notes means no reading was taken (for example, an aborted test).

Make and model normalisation

The raw DVSA data records make and model as free text entered by the tester, which contains inconsistent capitalisation, spacing, misspellings and, for a small number of makes, genuine duplicate entries (for example “VW” and “VOLKSWAGEN”, or “MERC” and “MERCEDES-BENZ”). These are merged using an explicit, manually reviewed alias list rather than automatic fuzzy matching, so that genuinely distinct historic marques (for example “ROVER”, which is not the same manufacturer as “LAND ROVER”) are not merged incorrectly. See the README for the full list.

For newer vehicles, the raw modelfield is often a full derivative string rather than a model name — for example “PUMA ST-LINE X MHEV” and “PUMA TITANIUM MHEV” are the same model, a Ford Puma, with different trim and engine badges attached. Left alone, these fragment a single real-world model into dozens of near-identical pages and — more importantly — corrupt the statistics: a fragment’s pass rate reflects only that trim, not the model as a whole, and a bare “Puma” page would end up dominated by whichever fragment happened to keep the plain name (in this dataset, the 1997–2002 Puma coupé) rather than the modern crossover.

Trim variants are collapsed into a parent model in two ways: a small number of makes with systematic alphanumeric codes get a dedicated rule (BMW’s “320D M SPORT MHEV AUTO” becomes “3 Series”, taking the model’s leading digit; Mercedes- Benz’s bare class letters such as “C” or “A” become “C-Class” / “A-Class”; Land Rover’s several spellings of “Range Rover” are normalised before dispatching to Sport / Evoque / Velar / the flagship). Every other make goes through a generic pass that strips trailing trim, engine, and gearbox words (a versioned list) and any trailing token containing a digit, stopping as soon as one token is left — so “TRANSIT CUSTOM 300LIMITD EBLUE” becomes “TRANSIT CUSTOM” without an explicit rule for that exact trim.

This means a deliberate, documented simplification for v1: a trailing performance or edition badge folds into its parent nameplate rather than getting its own page (Fiesta ST → Fiesta, Peugeot 208 GT → 208, Audi e-tron GT → e-tron), even where enthusiasts would consider it a distinct variant. Genuinely distinct nameplates that share a platform are kept separate rather than folded — Ford Puma vs Mustang Mach-E, VW Golf vs e-Golf vs ID.3, BMW 3 Series vs i4, Kia Niro vs EV6, Tesla Model 3 vs Model S/X/Y — because collapsing those would misrepresent genuinely different vehicles as one.

The full, versioned map of every raw model string that was collapsed or excluded is published alongside the aggregate data at data/normalisation/model-map.jsonin the site’s repository, so the exact effect of this normalisation on any given make is inspectable rather than opaque.

Minimum sample size

A make or model only gets a page on this site if it has at least 1,000 qualifying tests in the dataset. This avoids generating pages from small, noisy samples.

Within a model’s page, an individual vehicle-age band only shows a pass rate or average mileage if it has at least 30 qualifying tests. A model can clear the 1,000-test threshold overall while still having only a handful of tests in one age band (for example, very few near-new examples of an otherwise common older model) — showing a rate from a single-digit sample would be misleadingly precise, so those bands are marked as having too few tests rather than shown as a rate.