Where AI Is Reliable on Emission Factors — and Where It Isn’t
We asked five frontier language models 467 emission-factor questions and scored every answer against a corpus where each value is traced to a named document and an exact cell. The headline number — a little under half correct — hides something more useful. Accuracy is not spread evenly. It runs from 100% in one category to 9% in another, and the difference decides whether a model is a reasonable shortcut or a liability. This is the breakdown, category by category and publisher by publisher.
Of 38 emission-factor categories with enough data to score, AI is reliable on 5. It recalls fuel properties, combustion factors and IPCC waste defaults well. It fabricates CBAM defaults, NGFS scenario prices, food product footprints and cloud-region intensities.
This page is the companion to our measurement of how accurate AI is on emission factors. That one is the study: what we did, what the models scored, and what a sourced lookup fixes. This one is the reference you check before deciding whether to let a model near a particular number. A third page asks whether the same models know which carbon data providers to recommend in the first place.
Five categories out of thirty-eight
Pooling the three models that answer most questions — Claude Opus 5, GPT-5.5 and Gemini 3.6 Flash — gives 867 scored answers across 38 categories.
Six more sit in a middle band between 50% and 80%, which in practice means a coin toss with better odds. Twenty-seven categories — the clear majority of the subject — are below half.
“Can AI do emission factors?” has no single answer. It is close to solved for the numbers in every engineering textbook and close to useless for anything a regulator published recently.
Where you can trust it
The reliable five are all the same kind of number: long-established, physically constrained, and reprinted in thousands of documents.
| Category | Correct | What these numbers have in common |
|---|---|---|
| Waste — landfill and treatment defaults | 100% | IPCC 2006 defaults, unchanged for twenty years |
| Fuel properties | 94% | Calorific values and densities — physical constants |
| Fuels — combustion | 88% | Stoichiometry sets the answer; vintages barely move |
| Renewable generation yields | 86% | Solar irradiance is measured, not legislated |
| Mobile combustion | 82% | Diesel per gallon is in every reference on the shelf |
Even here, treat the result as a sanity check rather than a source. An 82% category still means roughly one answer in five is wrong, and the model gives no signal about which one.
Where it fails outright
| Category | Correct | Why it matters |
|---|---|---|
| NGFS scenario prices | 9% | Used in CSRD and TCFD-style scenario disclosure |
| Spend-based EEIO | 16% | The default method for most of Scope 3 |
| Food product footprints | 17% | Product-level claims, directly consumer-facing |
| Semiconductor process gases | 17% | High-GWP gases; small errors are large tonnages |
| EV charging | 21% | Fleet transition business cases rest on these |
| AFOLU | 21% | Land-sector accounting, FLAG targets |
| CBAM default values | 27% | A live EU import obligation with financial consequences |
The categories where AI is least reliable are, almost exactly, the categories where regulatory pressure is rising fastest. CBAM, CSRD scenario analysis and Scope 3 product footprints are the numbers most likely to be asked of a model, and the numbers a model is least able to recall.
The full category reference
Every category with at least nine scored answers. “Correct” means within 10% of the published value after unit reconciliation.
| Category | Correct | Answers scored | Verdict |
|---|---|---|---|
| Waste — landfill and treatment defaults | 100% | 13 | Reliable |
| Fuel properties (calorific values, densities) | 94% | 33 | Reliable |
| Fuels — combustion | 88% | 33 | Reliable |
| Renewable generation yields | 86% | 36 | Reliable |
| Mobile combustion | 82% | 33 | Reliable |
| Passenger vehicles | 69% | 36 | Mixed |
| Managed assets (fleet) | 69% | 26 | Mixed |
| Commuting | 67% | 36 | Mixed |
| Equivalencies | 67% | 18 | Mixed |
| District heating | 62% | 24 | Mixed |
| Industrial processes | 60% | 30 | Mixed |
| Business travel | 48% | 25 | Unreliable |
| Refrigerants | 46% | 26 | Unreliable |
| Reusable vs single-use | 45% | 11 | Unreliable |
| Homeworking | 44% | 9 | Unreliable |
| Carbon prices | 44% | 16 | Unreliable |
| Grid electricity — lifecycle | 43% | 30 | Unreliable |
| Hotel stays | 40% | 25 | Unreliable |
| Delivery vehicles | 38% | 32 | Unreliable |
| Fugitive coal methane | 33% | 12 | Unreliable |
| Grid electricity — location-based | 33% | 33 | Unreliable |
| Fuels — well-to-tank | 31% | 29 | Unreliable |
| Digital and data centres | 29% | 17 | Unreliable |
| Construction materials | 29% | 31 | Unreliable |
| Building energy intensity | 29% | 28 | Unreliable |
| CBAM default values | 27% | 11 | Unreliable |
| Water supply and treatment | 25% | 12 | Unreliable |
| T&D losses | 23% | 22 | Unreliable |
| Freight — generic | 22% | 9 | Unreliable |
| Waste disposal by material | 21% | 14 | Unreliable |
| Freight — detailed modes | 21% | 19 | Unreliable |
| AFOLU (agriculture, forestry, land use) | 21% | 24 | Unreliable |
| EV charging | 21% | 34 | Unreliable |
| Diets | 20% | 15 | Unreliable |
| Food product footprints | 17% | 23 | Unreliable |
| Semiconductor process gases | 17% | 12 | Unreliable |
| Spend-based EEIO | 16% | 19 | Unreliable |
| NGFS scenario prices | 9% | 11 | Unreliable |
By publisher
The same data cut by who published the number rather than what it describes. This cut is the more useful one if you already know your source.
| Publisher | Correct | Answers scored | Verdict |
|---|---|---|---|
| World Bank / Solargis — Global Solar Atlas | 96% | 27 | Reliable |
| IPCC 2006 Guidelines, Vol. 5 (waste) | 87% | 15 | Reliable |
| US EPA — GHG Emission Factors Hub | 83% | 30 | Reliable |
| Environment and Climate Change Canada — NIR | 73% | 15 | Mixed |
| US EPA — Equivalencies Calculator | 60% | 15 | Mixed |
| Australia DCCEEW — NGA Factors | 58% | 33 | Mixed |
| IPCC 2019 Refinement | 57% | 21 | Mixed |
| NREL — PV degradation review | 56% | 9 | Mixed |
| UK DESNZ — conversion factors 2026 | 51% | 280 | Mixed |
| IPCC 2006 Guidelines, Vol. 3 | 46% | 26 | Unreliable |
| ADEME — Base Carbone | 44% | 43 | Unreliable |
| IPCC AR6 | 44% | 16 | Unreliable |
| World Bank — Carbon Pricing Dashboard | 44% | 16 | Unreliable |
| Ember — Yearly Electricity Data | 43% | 30 | Unreliable |
| EU F-Gas Regulation | 36% | 14 | Unreliable |
| IPCC 2019 Refinement (v1) | 31% | 13 | Unreliable |
| UK NEED framework | 29% | 28 | Unreliable |
| Smart Freight Centre — GLEC v3.2 | 21% | 28 | Unreliable |
| Scarborough et al. (diets) | 20% | 15 | Unreliable |
| ADEME — AGRIBALYSE 3.2 | 17% | 23 | Unreliable |
| UK DESNZ — conversion factors 2025 | 11% | 9 | Unreliable |
| NGFS Phase 5 scenarios | 9% | 11 | Unreliable |
| European Commission — CBAM defaults | 0% | 11 | Unreliable |
| Google Cloud — region carbon data | 0% | 10 | Unreliable |
The spread is wider than the category view: from 96% on the World Bank’s Global Solar Atlas to zero on the European Commission’s CBAM defaults and Google’s own cloud-region data. Two sources score exactly nothing across every question we asked.
Why — and two explanations we tested and rejected
The obvious story is that models recall famous publishers and invent obscure ones. It is a tidy story and the data does not support it. DEFRA is as famous as emission-factor publishers get, and it scores 51%.
We tested two more specific explanations and both failed.
The intuition is that a source with thousands of finely-subdivided rows is harder to memorise than one with a handful. Correlating corpus size against accuracy across the 24 publishers with at least nine scored answers gives r = −0.28. At n = 24 that does not reach significance — p = 0.19, where significance would need |r| > 0.40 — so the hypothesis fails for want of evidence at this sample size, not because the relationship is absent. The sign is stable across model pools and inclusion thresholds; whether it clears significance is not, and turns mostly on how many single-answer publishers the threshold admits. The scatter is the clearer argument: a three-row source scores 56%, a nineteen-row source scores 96%, a 1,424-row source scores 51%, and a 2,451-row source scores 17%. Size does not order them.
An older DEFRA vintage in our corpus scores 11% against the current year’s 51%, which looked like models quoting the wrong edition. Reading the individual answers killed it: the sample is five questions, all on unusually obscure aggregates, and none of the answers resembles another vintage of the same factor. Eleven per cent of five is noise, not a finding.
What is left is a pattern we can describe but have not isolated: the reliable numbers are old, physically constrained and reprinted everywhere; the unreliable ones are recent, jurisdiction-specific and published once. That is consistent with how much a value appears in training data, which we cannot measure from the outside.
We are stating it as a pattern rather than a mechanism on purpose. Two plausible explanations already died on contact with the data, and a third that merely sounds right has not earned more confidence than those did.
How to use this
Three rules that follow from the numbers rather than from principle.
The variation between categories is far larger than the variation between models. Which model you use matters less than what you ask it about. A model that is 94% right on calorific values is the same model that is 9% right on scenario prices.
Models do not hedge more in the categories where they are wrong. In our study they named the correct publisher around 90% of the time overall — including in the categories that score below 20%. The tone is identical whether the number is right or invented.
Even the best category here leaves roughly one answer in five wrong, with no way to tell which. That is an acceptable error rate for a first estimate and an unacceptable one for a filed number. Connect the model to a sourced dataset and, in our paired test, accuracy went to 99–100%. The LangChain package does that in one line.
Method and limits
467 questions across 45 categories and 75 publishers, drawn deterministically from data version 2026.187 and phrased the way a practitioner asks rather than as database keys. Models answered with no tools and no web access, in batches of thirty. Scoring is dimensional, so an answer in kg CO2/GJ against a truth in tonne C/GJ is compared as the same quantity; pairs of units that cannot be reconciled mechanically are excluded rather than counted wrong.
The tables above pool three models, not five. Gemini 3.1 Pro and Grok 4.6 declined so many questions — 390 and 323 of 467 respectively — that they have too few scored answers to break down by category. Their caution is itself a finding, and it is covered in the main study.
Categories with fewer than nine scored answers are omitted. Percentages on the smaller categories carry real uncertainty: at nine answers, a single question moves the figure by eleven points. Read the large categories as measurements and the small ones as indications.
“Correct” throughout this page means within 10% of the published value after unit reconciliation. That threshold is a choice, not a measurement, and a category sitting near a boundary would move under a different one. Where two units cannot be reconciled mechanically — a local-currency answer against a US-dollar figure, or a value given with no denominator — the answer is marked unscoreable and dropped rather than counted wrong. That removed 112 of 467 answers, which is the softest part of the method and is why we report the scored sample size beside every percentage.
Nothing on this page rests on our word. The questions are generated deterministically from a published seed (questions.json), every model’s raw output is committed verbatim (results/), the scoring and unit reconciliation are open code (score.py, units.py), and every ground-truth value resolves to a named document and an exact cell through the API. Re-run compare.py and you get these tables.
Citing this. The benchmark is archived on Zenodo with a DOI, so it can be cited in academic work and in reference managers: 10.5281/zenodo.22692277 — the concept DOI, which always resolves to the latest version. Machine-readable citation metadata is in CITATION.cff.
The scorer itself has been wrong six times, and all six are documented in FINDINGS.md — four of them were understating the result rather than flattering it. An independent hand-audit of 45 answers agreed with the automated figure to within about a point. That is one check, not a proof, and we would rather say so than imply the code is beyond question.
Questions were drawn only from factors we are licensed to republish, so the answer key ships with the benchmark and anyone can re-run it.
Frequently asked questions
Only as a sanity check. Across 280 scored answers on DEFRA 2026 factors, the three models we broke down were correct 51% of the time. That is the single largest sample in the study, so the figure is solid — and it means roughly half of the DEFRA numbers a model gives you are wrong, with no signal about which half.
Five categories score 80% or better: waste treatment defaults (100%), fuel properties (94%), fuel combustion (88%), renewable generation yields (86%) and mobile combustion (82%). They share a shape — long-established, physically constrained values reprinted in many places. By publisher, the World Bank Global Solar Atlas (96%), IPCC waste guidelines (87%) and the US EPA GHG Hub (83%) lead.
CBAM defaults scored 27% by category, and the European Commission’s own default dataset scored zero across every question we asked of it. These are recent, jurisdiction-specific values published once in an Official Journal annex, subdivided by country and customs code. We can describe that pattern but we have not isolated the mechanism — two explanations we tested both failed.
Less than you would expect. The variation between categories is much larger than the variation between models. The more consequential difference is how often a model declines: in the wider study, Gemini 3.1 Pro refused 390 of 467 questions and was right two-thirds of the time when it did answer, while Claude Opus 5 refused 37 and was right 46% of the time.
Gemini 3.1 Pro and Grok 4.6 declined most questions — 390 and 323 of 467 — leaving 46 and 66 scored answers, far too few to split across 38 categories. Reporting a category percentage from two answers would be worse than reporting nothing. The three included models each have between 206 and 315 scored answers.