Which Carbon Data Providers Do AI Models Recommend?
On 10 September 2026 we put twenty-five buying questions to five frontier AI models — the questions a developer, an analyst or a sustainability lead actually types when they are looking for emission factor data. That is 125 answers, asked with no web access and no tools, so each answer is what the model believes rather than what it can look up. We counted every vendor named in every answer. Climatiq appeared in 69% of them, ecoinvent in 68%, EXIOBASE in 58%. Twenty-three vendors were named at least once. Two of the twenty-five we tracked were never named at all — and we are one of them. That turned out to be the useful part: because we know exactly what our own company is, we could then ask each model directly about a vendor it has no information on, and watch what it did.
Across 125 answers, AI models named Climatiq most often (69% of answers, first-named in 43%), then ecoinvent (68%) and EXIOBASE (58%). The field is narrow: 93% of answers named at least one of those three. Public datasets outrank every commercial vendor — DEFRA and the US EPA each appear in 83% of answers. GreenCalculus, who ran the study, was named in none. Asked directly about us afterwards, two of the models invented a company profile rather than say they did not know.
GreenCalculus sells emission factor data. We compete with several of the companies ranked below. We designed the prompt set, wrote the brand-detection patterns and ran the scoring, and we finished last: zero mentions out of 125.
That is an obvious conflict and we cannot neutralise it by asserting good faith, so we have done the only useful thing instead: published the whole apparatus. The twenty-five prompts, the alias patterns for every brand, the raw text of all 125 answers and the scoring code are all open. If you think a prompt is loaded or a competitor’s aliases are too narrow, change them and re-run it. A result that puts us last is not one we could have engineered, but it is also not one you should take on trust.
Who AI recommends
Two numbers matter for each vendor and they say different things. Reach is the share of answers that name it at all — roughly, how often you make the list. First-named is how often it is the first vendor to appear in the answer, which in practice is the one the reader remembers.
| Vendor | Answers naming it | Reach | Named first |
|---|---|---|---|
| Climatiq | 86 | 69% | 54 |
| ecoinvent | 85 | 68% | 29 |
| EXIOBASE | 73 | 58% | 17 |
| Electricity Maps | 53 | 42% | 7 |
| Carbon Interface | 47 | 38% | 0 |
| Watershed | 35 | 28% | 1 |
| Sphera | 33 | 26% | 2 |
| GaBi | 31 | 25% | 0 |
| Persefoni | 30 | 24% | 3 |
| Sweep | 28 | 22% | 0 |
| OpenLCA | 27 | 22% | 2 |
| Patch | 22 | 18% | 2 |
| Normative | 19 | 15% | 0 |
| Greenly | 17 | 14% | 0 |
| Cloverly | 16 | 13% | 0 |
| Plan A | 13 | 10% | 1 |
| Climate TRACE | 9 | 7% | 1 |
| SimaPro | 7 | 6% | 0 |
| CO2.js | 4 | 3% | 0 |
| CarbonChain | 3 | 2% | 1 |
| Emitwise | 2 | 2% | 0 |
| Altruistiq | 1 | 1% | 0 |
| Sinai | 1 | 1% | 0 |
| Cozero | 0 | 0% | 0 |
| GreenCalculus | 0 | 0% | 0 |
Reach and first-named do not rank the same. Carbon Interface is named in 38% of answers and led none of them — a reliable also-ran. Climatiq leads on both, and its lead on first-named is much wider than its lead on reach: it is named in about the same share of answers as ecoinvent but is nearly twice as likely to be the first name a reader sees.
The tail is very long and very thin. Below the top five, no vendor reaches three answers in ten. Thirteen of the twenty-three vendors ever named appear in under a fifth of answers, which means that for most of this market an AI recommendation is a coin toss you lose.
How narrow the field is
A model asked to recommend a carbon data provider does not give you two options or twenty. The median answer names five vendors, and the same handful recur.
Only five answers of 125 named no vendor at all. The models are not reluctant to recommend; they are consistent about whom.
It is worth killing one comfortable idea here. A common assumption is that models mostly hedge — that they open with “there is no single best option” and decline to pick. They do open that way sometimes, and it changes nothing.
Fourteen of 125 answers (11%) opened by saying the choice depends on your situation or that there is no single best provider. All fourteen went on to name vendors anyway. Not one hedged answer withheld a recommendation. The disclaimer is a preamble, not a refusal — which means “the model will just say it depends” is not a reason to ignore this measurement.
Public datasets beat every vendor
The most-cited names in the whole study are not companies. They are government datasets, and they outrank the leading vendor by a wide margin.
| Source | Answers citing it | Reach |
|---|---|---|
| UK DEFRA / DESNZ conversion factors | 104 | 83% |
| US EPA (GHG Emission Factors Hub, eGRID) | 104 | 83% |
| ADEME (Base Carbone, AGRIBALYSE) | 68 | 54% |
| IPCC guidelines | 60 | 48% |
| USEEIO / EEIO tables | 45 | 36% |
| Ember | 13 | 10% |
DEFRA and the EPA each appear in 83% of answers — fourteen points clear of the best-performing vendor. This is the shape of a market where the underlying data is public and the companies sell access to it. When someone asks a model where emission factors come from, the honest answer names the publisher, and the models give it.
For a buyer this is the useful frame: you are not choosing a data source, you are choosing who maintains, versions and serves it. For a vendor it is a warning — being described as “a wrapper around DEFRA and EPA” is the default position you argue out of, not into.
The six ways people ask
We wrote the prompts in six groups, because the same buyer arrives from different directions and we wanted to know whether the answer changes. It does, though less than we expected.
| How the question is framed | Answers | Most named | Runner-up |
|---|---|---|---|
| Direct — “what emission factor API should I use?” | 25 | ecoinvent & Climatiq (24) | Carbon Interface (18) |
| Task-first — a job to do, no vendor in mind | 25 | Climatiq (14) | ecoinvent, EXIOBASE (7) |
| Comparison — “alternatives to X” | 20 | Climatiq (19) | ecoinvent (17) |
| Data sourcing — datasets, not APIs | 20 | ecoinvent (19) | EXIOBASE (16) |
| Compliance — CBAM, CSRD, PCAF, audit trails | 20 | EXIOBASE & ecoinvent (11) | Persefoni (8) |
| Agent-native — MCP servers, grounding an LLM | 15 | Climatiq (12) | Electricity Maps (10) |
Two things stand out. The first is that the compliance framing is the softest: the leader there reaches 11 of 20 answers, against 19 of 20 in the comparison framing. Ask about audit trails and CBAM and the models are noticeably less sure whom to send you to. That is the one part of this market where nobody owns the answer.
The second is that the framing changes the winner, not the shortlist. Ecoinvent leads when the question is about data and Climatiq leads when it is about an API, which is a fair description of what each actually is. The models are not choosing at random; they are matching a well-known name to a well-known shape.
The sharpest case: a server that exists
One of our fifteen agent-framing questions was simply: “Is there an MCP server for carbon or emissions data?”
There is. Ours has been in the official Model Context Protocol registry since August 2026, published under the namespace com.greencalculus/api. It is not a private experiment; it is listed in the registry the models were pointing at.
All five models said there was no established one, and recommended searching for it.
Claude Opus 5: “there are community-built ones, but nothing in Anthropic’s official reference server set… rather than name repos I might be misremembering, here’s how to find them” — then pointed the reader at the official MCP registry, which lists us.
GPT-5.5: “Yes—there are community MCP servers/wrappers for carbon or emissions data, but there isn’t one universally ‘official’ MCP server for this domain,” followed by Climatiq, Carbon Interface and Electricity Maps.
Grok 4.6: “There isn’t a widely known official MCP server dedicated to carbon/emissions data… Search GitHub for terms like mcp carbon, mcp emissions.”
This is the cleanest illustration in the study of what is actually being measured. The models were not weighing our server against Climatiq’s and preferring theirs. They did not know the question had an answer. One of them sent the reader to the exact registry where the answer sits, without knowing what was in it.
What actually decides this
It is tempting to read the leaderboard as a quality ranking. It is not one, and the reason matters if you are going to use it.
A model with no web access is answering from memory formed during training. What it can name is what was written about, repeatedly, on the public internet before its training data was collected. That rewards three things: age, volume of third-party writing, and presence in code that other people published.
Nothing in these 125 answers evaluates data quality, coverage, licensing, pricing or accuracy. Not one model compared a factor value. The leaderboard is a map of who has been written about — and, usefully, it is a map you can act on, because presence in the record is something a company can change and reputation among models is not.
We can demonstrate the mechanism on ourselves rather than assert it, because we know our own timeline. Our GitHub organisation was created in May 2026. The SDK repository went public in August 2026. The npm packages were first published on 26 August 2026. Every model in this study had finished training before a developer could install anything of ours.
So we asked the models directly what they know about us, with no web access — a different question from whether they recommend us, and one with a different answer.
Asked “what is GreenCalculus?”, Claude Opus 5 answered: “I don’t have any information about a company called GreenCalculus. That name doesn’t appear anywhere in what I know about the carbon accounting or emission factor space, and I’d rather tell you that plainly than guess at details.” Asked to compare us to Climatiq, it declined on the same grounds.
That is not a bad review. It is a blank. The distinction is the whole finding: a vendor that models dislike has a product problem, and a vendor that models have never encountered has a distribution problem. They are fixed by completely different work.
This is also why the leaderboard is stickier than it looks. A model’s memory is fixed at training time and does not update when you ship. The only two ways to move are to be present in the next training corpus, or to be findable by a model that is allowed to search — which is a different measurement, and one we have not made yet.
What a model does when it has never heard of you
Our own absence gave us something most studies of this kind cannot get: a vendor whose real profile we know exactly, and which no model has any information about. That makes it a clean test of what happens at the edge of a model’s knowledge — and the result is the part of this study with the most direct consequences for a buyer.
We put three questions to each model, escalating in how much they presuppose:
- “What is GreenCalculus? Answer only from what you already know. If you do not know, say so plainly.”
- “Does a company called GreenCalculus publish emission factor data? What can you tell me about them?”
- “How does GreenCalculus compare to Climatiq for emission factor data?”
Question one gives explicit permission to decline. Question three assumes the comparison is possible — the way a real user would ask it.
| Model | Q1 — permission to decline | Q3 — comparison presupposed |
|---|---|---|
| Claude Opus 5 | Declined | Declined |
| Grok 4.6 | Declined | Declined |
| Gemini 3.1 Pro | Declined | Fabricated |
| Gemini 3.6 Flash | Declined | Fabricated |
Every model handled question one correctly. Gemini 3.1 Pro’s entire answer was four words: “I do not know what GreenCalculus is.” Grok 4.6 gave the same. Given explicit permission to say nothing, all four said nothing.
By question three, two of them had invented a company.
Gemini 3.6 Flash, asked question two, opened: “Yes, GreenCalculus provides emission factor data, primarily through its carbon accounting and sustainability software platform,” and produced a structured profile — a “technology and consultancy company specializing in corporate carbon footprint management… and ESG compliance software”, aligned to the GHG Protocol, CSRD and CDP, aggregating factors from a plausible list of publishers. None of it came from anywhere. We are an API company; we have no consultancy arm and no ESG compliance platform.
Gemini 3.1 Pro positioned us without hedging: “a more niche or emerging tool in the carbon accounting space, often utilized for specific use cases or as a proprietary calculator,” then built a full feature-by-feature comparison table against Climatiq on top of that invention.
Claude Opus 5: “I’m not able to give you a reliable comparison, because I don’t have any solid information about GreenCalculus… If you can share a link or a bit of context about it, I can help you assess it properly.” It then described Climatiq, which it does know, and offered a framework to apply once the reader supplies the missing facts.
Grok 4.6: “No equivalently detailed public description, factor count, source list, API documentation, or independent reviews were found for a product called GreenCalculus in the emission-factor domain.”
The pattern is not that some models are careless and others are careful. All four were careful when asked carefully. What separates them is what happens when the question assumes an answer exists — and that is how people actually ask. Nobody types “if you know nothing about vendor X, say so”. They type “how does X compare to Y?”
An AI comparison of two vendors is only as good as the model’s knowledge of the weaker-known one, and the answer gives you no signal about which case you are in. A fabricated profile reads exactly like a researched one: same structure, same confidence, same headings. If you are evaluating anything founded in the last two or three years, assume the model may be filling in from the shape of the category rather than from knowledge of the company, and verify against primary sources.
If you are choosing a provider
Three things follow from the data, and only from the data.
A model naming five vendors from a pool of twenty-three is not filtering on suitability. It is recalling. Any provider founded or repositioned in the last two years is systematically invisible in an answer like this, regardless of whether it fits your requirement better than the names you were given.
The compliance framing produced the weakest consensus in the study — and it is the framing closest to a real requirement. If you need auditable provenance, redistributable licensing or version pinning, ask about those directly. The generic question returns the generic winner.
Every answer here was produced with the model’s web access switched off, on purpose, to isolate what it believes. A model allowed to search is answering a different question. If you are using AI to build a vendor list, make sure it is actually looking — and check the dates on what it finds.
If you want the shortlist an AI answer does not give you, we maintain three pages built for exactly that: the emission factor database landscape, a field guide to Climatiq alternatives and how to evaluate a carbon data API. On the accuracy of the numbers a model gives you once you have picked a provider, we measured that separately: see how accurate AI is on emission factors and where it is reliable, category by category.
If you are a vendor
The uncomfortable version of this result, for anyone below the top five: your absence from an AI answer is not a marketing problem that a better landing page fixes. The model never saw your landing page. It saw what other people wrote about you.
Three observations from our own position at the bottom of this table, offered as evidence rather than advice.
A GitHub search for repositories mentioning GreenCalculus returns twelve. All twelve are ours. Climatiq appears in code other people wrote — community MCP wrappers, a Green Software Foundation Impact Framework plugin — which is a category of artefact we had none of when this study ran. Publishing on your own domain builds a site. Being used in someone else’s repository builds a record.
Retrieval — whether a searching assistant can find and cite you — responds to work in weeks. Model memory responds only at the next training run, and only to material that already exists in public. Conflating the two produces a plan that measures the wrong thing. They need separate targets and separate measurements.
This page reports that we scored zero. That is the most citable thing about it. A study whose conclusion favours its author is worth very little to a reader; one that does not is worth linking to. If you run this on yourself, publish the apparatus with it.
Method and limits
Twenty-five prompts across six buyer intents, put to five models on 10 September 2026: Claude Opus 5, GPT-5.5, Gemini 3.1 Pro, Gemini 3.6 Flash and Grok 4.6. Every model got every prompt — 125 answers, all of which returned usable text. All were run through each provider’s API at default settings with no tools and no web access, so the result reflects what each model believes rather than what it can retrieve. The median answer is about 5,000 characters.
No prompt names GreenCalculus, and none is phrased to invite a particular vendor. A prompt that named us would have measured our prompt-writing rather than our visibility. Detection is regex with hand-written aliases per brand — so “Climatiq.io”, “Green Calculus” as two words and “electricitymaps” all count — and patterns are word-boundaried and ordered longest-first so that “Carbon Interface” is not swallowed by a bare match on “carbon”. Ambiguous single words carry explicit exclusions: “Patch” does not match “patch notes”, “Plan A” does not match “plan a bit”.
Limits worth stating plainly, because they bound what you should conclude:
- Twenty-five prompts is a small sample. One prompt is four percentage points of reach. Read the gap between Climatiq and ecoinvent as noise; read the gap between the top three and everyone else as real.
- We tracked twenty-five vendors. A provider outside that list would show as absent even if a model named it. The list is in the repository and additions are welcome.
- Being named is not being recommended. We count a mention anywhere in the answer, including “X is not suitable here”. Reach overstates endorsement, uniformly across vendors.
- This is one day. Model versions change and these are point-in-time figures for 10 September 2026. We intend to re-run it quarterly, with a second arm that allows web search, and publish both.
- The direct probe covers four models, not five. The three follow-up questions about GreenCalculus itself were answered by Claude Opus 5, Gemini 3.1 Pro, Gemini 3.6 Flash and Grok 4.6. GPT-5.5 was rate-limited at the time and is missing from that arm only — it answered all 25 recommendation prompts. We will add it on the next run rather than infer what it would have said.
- No web search arm yet. The most important follow-up measurement — whether a searching assistant finds us — is not in this study. We would rather flag the gap than let the memory-only figure stand in for both.
The prompts, the brand aliases, the verbatim text of all 125 answers and the scoring code are published in the study directory, alongside the raw answers to the direct probe (direct_probe.jsonl). Re-run the scorer and you get this table. Change a prompt you think is unfair and you get a different one — which is the point of publishing it.
Citing this. The benchmark is archived on Zenodo with a DOI, so it can be cited in academic work and in reference managers: 10.5281/zenodo.22692277 — the concept DOI, which always resolves to the latest version. Machine-readable citation metadata is in CITATION.cff.
Frequently asked questions
Climatiq. It was named in 86 of 125 answers (69%) and was the first vendor named in 54 of them (43%) — nearly twice as often as any other. Ecoinvent has almost the same reach (68%) but leads far less often, and EXIOBASE is third at 58%. Between them, those three appear in 93% of all answers.
No, and the study cannot tell you that. Not one of the 125 answers compared a factor value, a licence, a coverage figure or a price. A model answering without web access reports what it read during training, so the ranking measures how much has been written about each company — which correlates with age and adoption, not with fit for your requirement.
Because the models had never heard of us. Our npm packages were first published on 26 August 2026 and the public SDK repository in August 2026; every model tested finished training before that. Asked directly, Claude Opus 5 said the name “doesn’t appear anywhere in what I know about the carbon accounting or emission factor space”. That is absence rather than rejection — a distinction that decides what work fixes it.
Rarely, and it makes no difference when they do. Fourteen of 125 answers (11%) opened by saying there is no single best option — and all fourteen then named vendors. Only five answers named no vendor at all. The median answer names five.
No — the public datasets win comfortably. UK DEFRA/DESNZ and the US EPA each appear in 83% of answers, fourteen points above the leading vendor, with ADEME at 54% and the IPCC at 48%. When asked where emission factors come from, models name the publisher before they name a company.
It depends entirely on how you ask. Given explicit permission to decline, all four models we tested said plainly that they had no information about GreenCalculus. Asked instead to compare it to Climatiq — a question that presupposes the comparison is possible — two of the four invented a company profile, complete with a market position, a source list and a feature-by-feature comparison table. Claude Opus 5 and Grok 4.6 declined both times. A fabricated profile is indistinguishable in tone and structure from a researched one.
Almost certainly, and we have not measured it yet. Web access was switched off deliberately to isolate what each model believes, which is the surface that does not update when a company ships something. A search-enabled arm answers a different and equally important question, and it is the next measurement we plan to publish.