Generative AI is not sustainable today, and it is not unsustainable either. It is unmeasured.
That is the honest one-line answer. The evidence for it is short:
- None of the three leading frontier labs has published an audited emissions report.
- Most data centre operators do not track their own water use.
- The per-query figures that circulate are either self-reported or measured independently under conditions the vendors do not disclose.
Whether AI is sustainable depends on which model, on which grid, in which basin, and against what alternative. The people saying it is fine and the people saying it is a disaster are mostly arguing about a dataset that does not exist yet.
I have spent much of 2026 auditing this question vendor by vendor and resource by resource, publishing what I found. This article is the synthesis. It brings the pieces together, gives you the comparison you probably came here for, and lays out what a serious organisation should do about it, including what will change in the next few months whether the labs like it or not.
Key takeaways
- The verdict is “unmeasured”, not “fine” or “catastrophic”. Anthropic, OpenAI and xAI have all reached IPO-stage without an audited emissions report. Google and Microsoft publish, and their totals are rising.
- Per-query, current models are cheap. A median Gemini prompt is around 0.24 Wh and 0.26 ml of water by Google’s own accounting; the most efficient independently measured Claude model leads its class. That is the small story.
- Infrastructure is the big story. Gas turbines at Colossus and Stargate, Scope 3 driving Microsoft’s emissions up 23% while its own operations fell 30%, and water concentrated in stressed basins.
- The comparison people want (Claude vs ChatGPT) has an answer, with caveats. Claude wins on measured per-query efficiency and now shares Colossus with xAI. OpenAI’s only per-query number is a CEO blog post. Neither discloses.
- The clock is now running. California SB253 requires Scope 1 and 2 reporting from November 2026. IPO filings put disclosure choices on the public record. Silence is about to get expensive.
Two pictures that are both true
Every serious discussion of AI’s footprint has to hold two pictures at once, and most public arguments fail by choosing one.
Picture one: per query, it is cheap and getting cheaper
- Google reports a median Gemini text prompt at 0.24 Wh, 0.03 g CO2e and 0.26 ml of water, and says energy per prompt fell 33-fold in the year to May 2025.
- The independent “How Hungry is AI?” benchmark placed Claude 3.7 Sonnet as the most efficient major model in its May 2025 measurement, at less than half the energy of OpenAI’s o3 on long-form work.
- Used with any care, the marginal cost of one exchange is genuinely small.
Picture two: in total, it is rising regardless
Per-query efficiency multiplied by explosive growth in query volume, on infrastructure built faster than clean power can be procured, gives a rising total however good each query gets.
- Microsoft: Scope 1 and 2 emissions down almost 30% since 2020. Total emissions up 23.4% over the same period, because Scope 3 from data centre construction swamped the operational gains.
- Google: total emissions up 51% on 2019.
The efficiency story is real. It is also the smaller of the two, a point made at length in the Anthropic audit.
The comparison: what each lab discloses and where it runs
This is the table most readers arrive looking for. It compresses three full audits into one view. Each row links to the detailed piece where the sourcing lives.
| Anthropic (Claude) | OpenAI (ChatGPT) | xAI (Grok) | Google (Gemini) | |
|---|---|---|---|---|
| Audited emissions report | None | None | None | Yes, third-party assured |
| Scope 1, 2, 3 published | No. Measurement underway with Watershed (Aug 2026) | No. Sustainability steering committee; hiring ESG reporting lead (Aug 2026) | No. Revised S-1 names water scarcity as a business risk, nothing on emissions | Yes. Total emissions up 51% vs 2019 |
| Public climate target | None | None | None | Net zero by 2030; 120% water replenishment |
| Per-query figure | None self-published. Independent: most efficient major model (May 2025) | 0.34 Wh and 0.000085 gal per average query, CEO blog post, no methodology | None. Not on the AI Energy Score leaderboard | 0.24 Wh, 0.03 gCO2e, 0.26 ml water per median prompt (own methodology, not third-party verified) |
| Infrastructure story | Took over all 300 MW of Colossus 1 (Memphis, May 2026); site is combined-cycle gas with a history of unpermitted turbines | Stargate: ~7 GW planned, over $400bn committed; Abilene runs partly on on-site gas turbines | Colossus 2 (Southaven): 33 unpermitted gas turbines, federal Clean Air Act suit, injunction pending; alleged 2,508 t NOx/yr | Owned fleet; reclaimed water at ~1/4 of campuses; per-facility water reporting |
| Climate action on record | Joined Frontier carbon-removal coalition (Jun 2026) | None. “Stargate Community” commitment on water and ecosystems, unquantified | None. Memphis wastewater plant paused (Apr 2026) | Extensive, including PPAs and removal purchases |
| Water disclosure | None | None | None | Per data centre, the most transparent of the hyperscalers |
Is Claude better for the environment than ChatGPT?
People ask this constantly, so it deserves a direct answer instead of a shrug.
| Verdict | Why | |
|---|---|---|
| Per-query efficiency | Claude | Only independent measurement available: Claude’s efficient models lead; OpenAI’s o3 was among the hungriest, at over double the most efficient competitor |
| Disclosure | Draw, and a poor one | Neither publishes audited emissions, a climate target or water figures. Anthropic joined Frontier; OpenAI has an unquantified community commitment. Both are now hiring the people who will eventually write the reports |
| Infrastructure | Draw, converging downward | Anthropic runs on Colossus 1, the Memphis site with a gas-turbine history. OpenAI builds Stargate on partly on-site gas generation |
Honest summary: Claude, marginally, on measured efficiency, and it is a distinction that matters far less than how much you use either tool and what you use it for. Per-query comparisons are worth having; they cannot carry the weight people put on them.
Energy: where the numbers actually come from
Three kinds of energy figures circulate, and they should not be mixed.
- Vendor self-reports. Google’s 0.24 Wh median prompt is the best-documented, with a published methodology that includes idle capacity, cooling overhead and water. Google itself notes it is a point-in-time figure and not third-party verified. Sam Altman’s 0.34 Wh for ChatGPT is a blog post.
- Independent benchmarks. The “How Hungry is AI?” study measures models under controlled conditions and is the reason we can say anything comparative at all. Its limitation is that it measures the model, not the data centre it runs in.
- Training runs. Grok 4 training consumed roughly 310 GWh, on the order of 154,000 tonnes of CO2 from that one run. Training is lumpy and public estimates are rough, but it is not negligible, and it recurs with every frontier model.
Below the vendor level, the lever you actually control is workflow design. In my own practice, four habits cut compute by roughly half on sustained work: caching repeated context, loading tools only when needed, smaller models for routine tasks, and no AI image generation for structured visuals. The detail is in how Claude Code skills cut AI energy use. It is the one part of the footprint where a user’s decisions change the number today.
Water: mostly in the electricity, and mostly unmeasured
Water is the resource where public understanding lags furthest behind the research, so I gave it a full article. The short version has three parts.
- Around three-quarters of AI’s water is consumed at power stations generating the electricity, not inside the data centre. Water and energy are the same question in different units.
- National averages cannot settle it. Water stress is a basin-level, monthly property; 2026 research shows fixed national factors misestimate stress-adjusted water by up to 75% low and 999% high depending on region.
- The measurement gap is worse than for carbon. Fewer than a third of operators measured water in 2021, two-fifths in the UK today, and Amazon still publishes nothing.
For the UK specifically: 84% of proposed water-intensive data centres sit in areas already water-stressed or projected to be by 2040, and the water companies’ 2050 deficit plans do not account for them at all. That is a different problem from the US one, and the sceptical case built on US national totals does not transfer.
Why the labs stay quiet, and why that is about to change
Bloomberg put it plainly this week: SpaceX, OpenAI and Anthropic are heading toward public markets with almost no emissions data on the table, in what it called a post-ESG Wall Street. Investors are not demanding it, so it is not being volunteered. That is a rational reading of the market as it stood in 2025.
Three things are now moving against that position.
- California SB253. From November 2026, companies with over $1 billion in California revenue must report Scope 1 and 2 emissions. OpenAI has said it is coordinating with data centre partners ahead of the deadline. This is the first mandatory trigger, and it catches all three labs.
- The IPO filings. SpaceX’s S-1 is public and its revised version now names water scarcity as a business risk while saying nothing about emissions. Anthropic and OpenAI have filed confidentially. Whatever they choose to disclose becomes permanent public record.
- The hiring. Anthropic is working with Watershed to measure its footprint and has recruited for sustainability reporting and infrastructure energy accounting. OpenAI is hiring a head of ESG reporting. Companies do not build reporting teams for reports they intend never to publish.
The realistic expectation is that the first real disclosures arrive within a year, driven by regulation and listing rather than conviction. When they do, the per-query efficiency story will still be true, and the totals will be large. Both pictures, again.
Where I land
I use these tools every working day and I build products on them, so I have every incentive to conclude that they are fine. I have tried to argue the other side, and I end up in the same place across all four audits.
Sustainability is not a property a model has. It is a property of a whole system: model, data centre, grid, water basin, growth rate, and what the work displaces. Right now most of that system is unmeasured by the people best placed to measure it.
So the defensible position for anyone using generative AI seriously is a set of behaviours, not a verdict:
- Use less.
- Pick the smallest model that does the job.
- Cache and reuse repeated context.
- Ask vendors where the workload runs and on what grid, and treat a non-answer as an answer.
- Keep the pressure on for disclosure. The gap between the optimistic and pessimistic readings of AI’s footprint is almost entirely made of data that companies could publish tomorrow.
Five questions to ask any AI vendor
The same five I close each vendor audit with, because they have not changed and no lab yet answers all of them.
- Do you publish audited Scope 1, 2 and 3 emissions, and where?
- Where do the workloads I pay for physically run, and what is the grid mix there?
- What is your water consumption per facility, and how much of it is potable?
- What is your per-query energy figure, who measured it, and under what conditions?
- What have you committed to, by when, and who verifies it?
Inside an organisation, these belong in the governance perimeter alongside the confidentiality and accuracy rules, which is how I structure them in AI training for leadership teams.
Frequently asked questions
Is generative AI bad for the environment?
Per query, the impact is small and shrinking. In total, it is rising and largely unmeasured by the labs themselves. Both are true, and which one matters depends on whether you are deciding how to use a tool or how to regulate an industry.
Which AI model is the most sustainable?
On independent per-query measurement, Claude’s efficient models led the field as of May 2025. Google publishes the most complete self-reported figures. Neither fact settles the question, because the data centre and grid behind the model can outweigh the model’s own efficiency.
Does Anthropic publish an environmental report?
Not as of August 2026. It has joined the Frontier carbon-removal coalition and is measuring its footprint with Watershed, and it faces California SB253 reporting from November 2026. Full detail in the Anthropic audit.
How much energy does one AI prompt use?
Google reports 0.24 Wh for a median Gemini text prompt; OpenAI’s CEO cited 0.34 Wh for an average ChatGPT query. Reasoning models and long inputs use several times more. Roughly, a short exchange costs less than a few seconds of television; a long reasoning task with a large model costs meaningfully more.
What can I actually do about it?
Reduce demand before anything else. Smaller models for routine tasks, cached context on repeated work, no AI image generation where a chart or a stock photo does the job. Then ask your vendors the five questions above and prefer the ones who can answer.
