RoadVitals

Methodology

How the Vehicle Issue Index is calculated, what it can and cannot tell you, and where every number on this site comes from.

The Vehicle Issue Index is our own informational metric, not an official safety rating and not a reliability prediction.

What the Vehicle Issue Index is

In one sentence: the Vehicle Issue Index is a 0–100 summary of how much reported-issue activity a model year shows relative to comparable vehicles of the same body class and the same model year. Higher means more reported activity. That is all it means.

The Vehicle Issue Index is calculated by us from public NHTSA data. It summarises how much reported-issue activity a model year has relative to comparable vehicles of the same type and year. It is not an official safety rating, not a reliability prediction, and it says nothing about whether any individual vehicle will develop a problem.

Three things it is not

  • Not a safety rating. NHTSA’s 5-Star Safety Ratings are official U.S. government crash-test results. Ours is an activity summary we compute ourselves from public records. The two measure different things, and we never present one as the other.
  • Not a reliability prediction. It describes reports that have already been filed. It does not forecast breakdowns, repair bills or ownership costs, and we publish no estimate of any of those because we hold no data that would support one.
  • Not a statement about any individual vehicle. A score describes a model year in aggregate. Two vehicles of the same year can have entirely different histories, and nothing here can tell you whether yours will develop a fault.

The 6 inputs and their weights

Weights sum to 100%. The volume-independent column marks inputs that are rates or shapes — measures that do not grow simply because more of a vehicle was sold.

Inputs to the Vehicle Issue Index and their weights
InputWeightWhat it measuresVolume-independent
Complaint volume30%Where this vehicle ranks for total owner-reported complaints among comparable vehicles of the same model year.No
Severity mix20%The share of its complaints that report a crash, fire, injury or death.Yes
Recent activity15%Whether complaints are arriving faster than expected for a vehicle of this age.Yes
Recall activity15%Where it ranks for number of safety recalls among comparable vehicles of the same model year.No
Investigations10%How many NHTSA safety investigations name this vehicle.Yes
Issue concentration10%Whether complaints cluster on one system rather than being spread across many.Yes

How each input is scored

  • Complaint volume is a percentile rank inside the peer group, not a raw count, so it answers “more or fewer reports than comparable vehicles of the same year” rather than “a big number”.
  • Severity mix is the share of a vehicle’s complaints whose submitter recorded a crash, fire, injury or death. The scale is calibrated so that a 12% share sits at its midpoint and a 40% share at its practical ceiling. Those two figures are settings we chose, not measurements: 12% is set near the share carrying one of those flags across the complaint records we hold, so that a typical vehicle lands mid-scale rather than at either end.
  • Recent activity compares the share of complaints filed in the trailing 365 days against what a vehicle of that age would be expected to show (1 ÷ age in years). Being new is therefore not by itself unusual; being unusually active for its age is.
  • Recall activity is a percentile rank of recall count inside the same peer group.
  • Investigations use a step function rather than a percentile: most vehicles have zero, and a percentile over a mostly-zero distribution is noise. One investigation scores 45, two 65, three 80, and each further one adds 5.
  • Issue concentration is a normalised Herfindahl–Hirschman index over component categories — 0 when complaints are spread evenly, 100 when they all name one system. It is normalised against the lowest value achievable for the number of categories present, so a vehicle with few categories is not automatically scored as concentrated.

Age bias, and how peer groups handle it

A 2012 vehicle has had more than a decade to accumulate complaints; a 2025 vehicle has had months. Comparing their raw totals says nothing about either one. So no vehicle is ever scored against the corpus as a whole: every percentile is computed inside a peer group of the same model year, subdivided by body class so that a compact car is not ranked against a heavy-duty pickup.

Where the source data records no body class for a vehicle, it is grouped with the other vehicles of its model year that also have none, and the peer-group label on its page reads “All vehicles” for that year rather than a class name — so you can see when a comparison is coarser than usual.

Popularity bias, which we cannot fully solve

This is the real limitation of the metric, and we would rather state it here than have you discover it yourself. A truck that sold 700,000 units will out-complain a sports car that sold 8,000 regardless of how either was built. Correcting for that requires an exposure denominator — units sold or registered, per model year.

There is no free, reliable, per-model-year US sales or registration dataset available to us, and we refuse to invent an exposure denominator. Manufacturer sales figures are published irregularly, usually at model rather than model-year granularity, and often not at all. Estimating the denominator would make every score look more rigorous while resting on a number we made up, so we do not do it.

What we do instead is mitigation, not a fix: 4 of the 6 inputs — carrying 55% of the total weight — are rates or shapes rather than counts, and the count-sensitive remainder is percentile-ranked inside a body-class peer group so that at least pickups are compared with pickups. A high-selling vehicle still tends to score higher than a rare one with an identical defect profile. Higher-selling and older vehicles accumulate more reports. Counts are not failure rates and are not directly comparable between vehicles that sold in very different numbers.

If we ever obtain exposure data we trust, complaint volume becomes a rate and the formula version below bumps.

Confidence, and when we show no number

A score computed from four complaints against a peer group of three is arithmetic, not evidence. Below the medium threshold we publish nothing at all rather than a number with a caveat attached, because a printed figure gets used regardless of what the caveat says.

Confidence thresholds for publishing an Issue Index
ConfidenceRequiresWhat we show
High30+ complaints and a peer group of 20+ vehiclesScore published.
Medium10+ complaints and a peer group of 8+ vehiclesScore published, labelled so that small differences are read as noise.
LowAnything below the medium rowNo score at all. The page says there is not enough data.

Reading a score

Band labels describe activity relative to the peer group and nothing else. There is no “good” or “bad” band, and a low score is not a recommendation.

Vehicle Issue Index bands
ScoreBandMeaning
0.024.9Below average activityLess reported issue activity than most comparable vehicles of this model year.
25.044.9Average activityReported issue activity in line with comparable vehicles of this model year.
45.069.9Above average activityMore reported issue activity than most comparable vehicles of this model year.
70.0100.0High activitySubstantially more reported issue activity than comparable vehicles of this model year.

How often this updates

Everything runs on a daily cycle, in UTC, because the upstream files are regenerated once a day:

  • Daily — NHTSA complaints, recalls, investigations and manufacturer communications are re-imported from the ODI flat files.
  • Weekly — EPA/DOE fuel economy and NHTSA 5-Star Safety Ratings. Both change on the order of weeks, and the ratings have no bulk export, so we query that API politely rather than often.
  • Daily, after the imports — per-vehicle aggregates and indexability are recomputed, then the Vehicle Issue Index is recomputed for every model year that clears the thresholds, then the day’s snapshot is written.

The snapshots are the part that cannot be rebuilt. NHTSA publishes only the current cumulative dataset — it has no concept of what a vehicle looked like three months ago. Every trend chart and every “complaints in the last 30 days” figure on this site is read from history we recorded ourselves, day by day. A missed day is gone permanently, so we monitor for it.

Formula version 1.0.0

Every computed score is stored with the version of the formula that produced it. The version is semantic: a major bump means the shape of the metric changed and scores are not comparable across versions, while a minor bump means tuning within the same shape. Any change to a weight or a sub-score definition bumps it, so we always know which rows need recomputing and you always know which rule produced the number in front of you.

What the source data cannot tell you

The metric is only as good as its inputs, and the inputs have real limits. These apply to the raw records on every vehicle page too, not only to the index.

  • Complaints are self-reported and unverified. Complaints are reports submitted by owners and drivers to NHTSA. They are not verified and do not establish that a defect exists. Reporting is also driven by attention: coverage of an issue produces more reports of that issue.
  • Recalls apply to VINs, not to model years. A campaign listed against a model year frequently covers only vehicles built in a particular window or at a particular plant. Our count is of campaigns that name the model year, which is not the same as the number affecting any given car.
  • Manufacturer communications are published only as summaries. NHTSA publishes the existence and a short summary of the bulletins manufacturers send their dealers, not the full documents, so that is all we can show. A bulletin sent by a manufacturer to its dealers. Not a recall, and repairs are not necessarily free.
  • Fuel economy joins imperfectly. EPA names vehicles differently from NHTSA and bakes drivetrain qualifiers into its model strings. We link an EPA configuration to a model year only on an exact normalised match, so fuel economy is simply absent on some vehicles rather than guessed at.
  • Crash-test ratings are not universal. NHTSA does not test every vehicle, ratings apply to the configuration tested, and the programme changed methodology in 2011, so earlier stars are not comparable with later ones.

Recalls apply to specific vehicles, not to every vehicle of a model year. Check your VIN with NHTSA or your manufacturer's dealer to confirm whether a recall affects your vehicle.

Check a VIN on NHTSA.gov

Corrections

If a page here contradicts its source record, tell us and we will fix it — email [email protected] with the URL and the campaign or ODI number. If the government record is itself wrong, we cannot silently edit our copy of it; that correction has to happen at the source, and we will point you at the right place. See data sources for what we take from each publisher, and about for how we work.

Data sources