Privacy Policy updated — version 1.7, 9 September 2026. A new section describes our Google Sheets add-on: what leaves your spreadsheet (only the factor names and search terms you enter) and what does not. Nothing new is collected, and no new processor is involved. Also recently: directory enquiries go through a form — we pass your message on and keep no copy — and anyone named in a listing can ask us to remove their details. Read the policy.

  1. Home
  2. Guides
  3. Reference
  4. Where AI Is Reliable on Emission Factors — and Where It Isn’t
Last reviewed September 2026
Authored by Jeremiah Say

Founder and Lead Systems Architect of GreenCalculus. Translates GHG Protocol methodology into high-precision JavaScript calculation engines. Architect of the MasterBrain data layer covering 16,000+ sourced emission factors, aligned with IPCC AR6 and the GHG Protocol Corporate Standard.

Full profile →

Verified by GreenCalculus Engineering

Automated verification pipeline that audits every page against its underlying calculation code, source documents, and MasterBrain data layer. Traces every figure cell-by-cell to its named source workbook, enforces cell-by-cell provenance attribution on every emission factor, and cross-checks methodology prose against the data layer to catch stated-vs-actual discrepancies before publication.

Governance & verification pipeline →

Where AI Is Reliable on Emission Factors — and Where It Isn’t

We asked five frontier language models 467 emission-factor questions and scored every answer against a corpus where each value is traced to a named document and an exact cell. The headline number — a little under half correct — hides something more useful. Accuracy is not spread evenly. It runs from 100% in one category to 9% in another, and the difference decides whether a model is a reasonable shortcut or a liability. This is the breakdown, category by category and publisher by publisher.

Quick Answer

Of 38 emission-factor categories with enough data to score, AI is reliable on 5. It recalls fuel properties, combustion factors and IPCC waste defaults well. It fabricates CBAM defaults, NGFS scenario prices, food product footprints and cloud-region intensities.

This page is the companion to our measurement of how accurate AI is on emission factors. That one is the study: what we did, what the models scored, and what a sourced lookup fixes. This one is the reference you check before deciding whether to let a model near a particular number. A third page asks whether the same models know which carbon data providers to recommend in the first place.

Five categories out of thirty-eight

Pooling the three models that answer most questions — Claude Opus 5, GPT-5.5 and Gemini 3.6 Flash — gives 867 scored answers across 38 categories.

5 of 38 categories where AI answers correctly at least 80% of the time 27 of 38 are below 50%

Six more sit in a middle band between 50% and 80%, which in practice means a coin toss with better odds. Twenty-seven categories — the clear majority of the subject — are below half.

Key Point

“Can AI do emission factors?” has no single answer. It is close to solved for the numbers in every engineering textbook and close to useless for anything a regulator published recently.

Where you can trust it

The reliable five are all the same kind of number: long-established, physically constrained, and reprinted in thousands of documents.

The categories AI recalls reliably
CategoryCorrectWhat these numbers have in common
Waste — landfill and treatment defaults100%IPCC 2006 defaults, unchanged for twenty years
Fuel properties94%Calorific values and densities — physical constants
Fuels — combustion88%Stoichiometry sets the answer; vintages barely move
Renewable generation yields86%Solar irradiance is measured, not legislated
Mobile combustion82%Diesel per gallon is in every reference on the shelf

Even here, treat the result as a sanity check rather than a source. An 82% category still means roughly one answer in five is wrong, and the model gives no signal about which one.

Where it fails outright

The categories AI should not be asked about
CategoryCorrectWhy it matters
NGFS scenario prices9%Used in CSRD and TCFD-style scenario disclosure
Spend-based EEIO16%The default method for most of Scope 3
Food product footprints17%Product-level claims, directly consumer-facing
Semiconductor process gases17%High-GWP gases; small errors are large tonnages
EV charging21%Fleet transition business cases rest on these
AFOLU21%Land-sector accounting, FLAG targets
CBAM default values27%A live EU import obligation with financial consequences
The uncomfortable overlap

The categories where AI is least reliable are, almost exactly, the categories where regulatory pressure is rising fastest. CBAM, CSRD scenario analysis and Scope 3 product footprints are the numbers most likely to be asked of a model, and the numbers a model is least able to recall.

The full category reference

Every category with at least nine scored answers. “Correct” means within 10% of the published value after unit reconciliation.

AI accuracy by factor category, three models pooled
CategoryCorrectAnswers scoredVerdict
Waste — landfill and treatment defaults100%13Reliable
Fuel properties (calorific values, densities)94%33Reliable
Fuels — combustion88%33Reliable
Renewable generation yields86%36Reliable
Mobile combustion82%33Reliable
Passenger vehicles69%36Mixed
Managed assets (fleet)69%26Mixed
Commuting67%36Mixed
Equivalencies67%18Mixed
District heating62%24Mixed
Industrial processes60%30Mixed
Business travel48%25Unreliable
Refrigerants46%26Unreliable
Reusable vs single-use45%11Unreliable
Homeworking44%9Unreliable
Carbon prices44%16Unreliable
Grid electricity — lifecycle43%30Unreliable
Hotel stays40%25Unreliable
Delivery vehicles38%32Unreliable
Fugitive coal methane33%12Unreliable
Grid electricity — location-based33%33Unreliable
Fuels — well-to-tank31%29Unreliable
Digital and data centres29%17Unreliable
Construction materials29%31Unreliable
Building energy intensity29%28Unreliable
CBAM default values27%11Unreliable
Water supply and treatment25%12Unreliable
T&D losses23%22Unreliable
Freight — generic22%9Unreliable
Waste disposal by material21%14Unreliable
Freight — detailed modes21%19Unreliable
AFOLU (agriculture, forestry, land use)21%24Unreliable
EV charging21%34Unreliable
Diets20%15Unreliable
Food product footprints17%23Unreliable
Semiconductor process gases17%12Unreliable
Spend-based EEIO16%19Unreliable
NGFS scenario prices9%11Unreliable

By publisher

The same data cut by who published the number rather than what it describes. This cut is the more useful one if you already know your source.

AI accuracy by publisher, three models pooled
PublisherCorrectAnswers scoredVerdict
World Bank / Solargis — Global Solar Atlas96%27Reliable
IPCC 2006 Guidelines, Vol. 5 (waste)87%15Reliable
US EPA — GHG Emission Factors Hub83%30Reliable
Environment and Climate Change Canada — NIR73%15Mixed
US EPA — Equivalencies Calculator60%15Mixed
Australia DCCEEW — NGA Factors58%33Mixed
IPCC 2019 Refinement57%21Mixed
NREL — PV degradation review56%9Mixed
UK DESNZ — conversion factors 202651%280Mixed
IPCC 2006 Guidelines, Vol. 346%26Unreliable
ADEME — Base Carbone44%43Unreliable
IPCC AR644%16Unreliable
World Bank — Carbon Pricing Dashboard44%16Unreliable
Ember — Yearly Electricity Data43%30Unreliable
EU F-Gas Regulation36%14Unreliable
IPCC 2019 Refinement (v1)31%13Unreliable
UK NEED framework29%28Unreliable
Smart Freight Centre — GLEC v3.221%28Unreliable
Scarborough et al. (diets)20%15Unreliable
ADEME — AGRIBALYSE 3.217%23Unreliable
UK DESNZ — conversion factors 202511%9Unreliable
NGFS Phase 5 scenarios9%11Unreliable
European Commission — CBAM defaults0%11Unreliable
Google Cloud — region carbon data0%10Unreliable

The spread is wider than the category view: from 96% on the World Bank’s Global Solar Atlas to zero on the European Commission’s CBAM defaults and Google’s own cloud-region data. Two sources score exactly nothing across every question we asked.

Why — and two explanations we tested and rejected

The obvious story is that models recall famous publishers and invent obscure ones. It is a tidy story and the data does not support it. DEFRA is as famous as emission-factor publishers get, and it scores 51%.

We tested two more specific explanations and both failed.

Rejected: bigger sources are harder

The intuition is that a source with thousands of finely-subdivided rows is harder to memorise than one with a handful. Correlating corpus size against accuracy across the 24 publishers with at least nine scored answers gives r = −0.28. At n = 24 that does not reach significance — p = 0.19, where significance would need |r| > 0.40 — so the hypothesis fails for want of evidence at this sample size, not because the relationship is absent. The sign is stable across model pools and inclusion thresholds; whether it clears significance is not, and turns mostly on how many single-answer publishers the threshold admits. The scatter is the clearer argument: a three-row source scores 56%, a nineteen-row source scores 96%, a 1,424-row source scores 51%, and a 2,451-row source scores 17%. Size does not order them.

Rejected: models confuse vintages

An older DEFRA vintage in our corpus scores 11% against the current year’s 51%, which looked like models quoting the wrong edition. Reading the individual answers killed it: the sample is five questions, all on unusually obscure aggregates, and none of the answers resembles another vintage of the same factor. Eleven per cent of five is noise, not a finding.

What is left is a pattern we can describe but have not isolated: the reliable numbers are old, physically constrained and reprinted everywhere; the unreliable ones are recent, jurisdiction-specific and published once. That is consistent with how much a value appears in training data, which we cannot measure from the outside.

We are stating it as a pattern rather than a mechanism on purpose. Two plausible explanations already died on contact with the data, and a third that merely sounds right has not earned more confidence than those did.

How to use this

Three rules that follow from the numbers rather than from principle.

1. Use the category, not the model, to decide

The variation between categories is far larger than the variation between models. Which model you use matters less than what you ask it about. A model that is 94% right on calorific values is the same model that is 9% right on scenario prices.

2. A confident answer carries no signal

Models do not hedge more in the categories where they are wrong. In our study they named the correct publisher around 90% of the time overall — including in the categories that score below 20%. The tone is identical whether the number is right or invented.

3. Anything going into a disclosure needs a source, not a recall

Even the best category here leaves roughly one answer in five wrong, with no way to tell which. That is an acceptable error rate for a first estimate and an unacceptable one for a filed number. Connect the model to a sourced dataset and, in our paired test, accuracy went to 99–100%. The LangChain package does that in one line.

Method and limits

467 questions across 45 categories and 75 publishers, drawn deterministically from data version 2026.187 and phrased the way a practitioner asks rather than as database keys. Models answered with no tools and no web access, in batches of thirty. Scoring is dimensional, so an answer in kg CO2/GJ against a truth in tonne C/GJ is compared as the same quantity; pairs of units that cannot be reconciled mechanically are excluded rather than counted wrong.

The tables above pool three models, not five. Gemini 3.1 Pro and Grok 4.6 declined so many questions — 390 and 323 of 467 respectively — that they have too few scored answers to break down by category. Their caution is itself a finding, and it is covered in the main study.

Categories with fewer than nine scored answers are omitted. Percentages on the smaller categories carry real uncertainty: at nine answers, a single question moves the figure by eleven points. Read the large categories as measurements and the small ones as indications.

“Correct” throughout this page means within 10% of the published value after unit reconciliation. That threshold is a choice, not a measurement, and a category sitting near a boundary would move under a different one. Where two units cannot be reconciled mechanically — a local-currency answer against a US-dollar figure, or a value given with no denominator — the answer is marked unscoreable and dropped rather than counted wrong. That removed 112 of 467 answers, which is the softest part of the method and is why we report the scored sample size beside every percentage.

Every percentage here can be reconstructed

Nothing on this page rests on our word. The questions are generated deterministically from a published seed (questions.json), every model’s raw output is committed verbatim (results/), the scoring and unit reconciliation are open code (score.py, units.py), and every ground-truth value resolves to a named document and an exact cell through the API. Re-run compare.py and you get these tables.

Citing this. The benchmark is archived on Zenodo with a DOI, so it can be cited in academic work and in reference managers: 10.5281/zenodo.22692277 — the concept DOI, which always resolves to the latest version. Machine-readable citation metadata is in CITATION.cff.

The scorer itself has been wrong six times, and all six are documented in FINDINGS.md — four of them were understating the result rather than flattering it. An independent hand-audit of 45 answers agreed with the automated figure to within about a point. That is one check, not a proof, and we would rather say so than imply the code is beyond question.

Questions were drawn only from factors we are licensed to republish, so the answer key ships with the benchmark and anyone can re-run it.

Where AI is reliable on emission factors, broken down by category and publisher
Save to Pinterest Download · 1000×1500 JPG

Frequently asked questions

Only as a sanity check. Across 280 scored answers on DEFRA 2026 factors, the three models we broke down were correct 51% of the time. That is the single largest sample in the study, so the figure is solid — and it means roughly half of the DEFRA numbers a model gives you are wrong, with no signal about which half.

Five categories score 80% or better: waste treatment defaults (100%), fuel properties (94%), fuel combustion (88%), renewable generation yields (86%) and mobile combustion (82%). They share a shape — long-established, physically constrained values reprinted in many places. By publisher, the World Bank Global Solar Atlas (96%), IPCC waste guidelines (87%) and the US EPA GHG Hub (83%) lead.

CBAM defaults scored 27% by category, and the European Commission’s own default dataset scored zero across every question we asked of it. These are recent, jurisdiction-specific values published once in an Official Journal annex, subdivided by country and customs code. We can describe that pattern but we have not isolated the mechanism — two explanations we tested both failed.

Less than you would expect. The variation between categories is much larger than the variation between models. The more consequential difference is how often a model declines: in the wider study, Gemini 3.1 Pro refused 390 of 467 questions and was right two-thirds of the time when it did answer, while Claude Opus 5 refused 37 and was right 46% of the time.

Gemini 3.1 Pro and Grok 4.6 declined most questions — 390 and 323 of 467 — leaving 46 and 66 scored answers, far too few to split across 38 categories. Reporting a category percentage from two answers would be worse than reporting nothing. The three included models each have between 206 and 315 scored answers.

Scroll to Top