Model ranking
Which AI model uses the most water?
Estimated intensity per 1,000 tokens for the models people actually name, at the Mid scenario (cooling + the water behind electricity). Highest measured intensity first.
Class assignments, benchmark anchors and official overrides are listed in full on the methodology page.
mL / 1K tokens · Mid scenario
- 01
ChatGPT (thinking / o3)
reasoning
7 Wh · 2.5 g CO₂e
35 mL
- 02
DeepSeek-R1
reasoning
7 Wh · 2.5 g CO₂e
35 mL
- 03
ChatGPT (standard)
frontier
2 Wh · 0.7 g CO₂e
10 mL
- 04
Claude Opus
frontier
2 Wh · 0.7 g CO₂e
10 mL
- 05
Gemini Pro
frontier
2 Wh · 0.7 g CO₂e
10 mL
- 06
Grok
frontier
2 Wh · 0.7 g CO₂e
10 mL
- 07
ChatGPT (mini)
standard
0.8 Wh · 0.28 g CO₂e
4 mL
- 08
Claude Sonnet
standard
0.8 Wh · 0.28 g CO₂e
4 mL
- 09
Mistral Large
standard
0.8 Wh · 0.28 g CO₂e
4 mL
- 10
Llama (70B+)
standard
0.8 Wh · 0.28 g CO₂e
4 mL
- 11
Gemini (app)
standard
0.34 Wh · 0.12 g CO₂e
1.7 mL
- 12
ChatGPT (nano)
efficient
0.3 Wh · 0.11 g CO₂e
1.5 mL
- 13
Claude Haiku
efficient
0.3 Wh · 0.11 g CO₂e
1.5 mL
- 14
Gemini Flash
efficient
0.3 Wh · 0.11 g CO₂e
1.5 mL
- 15
Llama (small)
efficient
0.3 Wh · 0.11 g CO₂e
1.5 mL
Mid scenario, mid benchmark values. Reasoning models rank highest because they generate long hidden chains of thought for every visible answer. Image and video models are excluded here: their unit is per generation, not per token.
Energy sets the order
Water and carbon here are energy times a fixed Mid factor, so all three tabs rank models the same way. Only the scale changes.
Classes, not guesses
Few models publish a per-token figure. Each one sits in a benchmarked class range, with an official number used instead wherever one exists.
Reasoning costs more
Reasoning models write long hidden chains of thought before they answer, so the same visible reply takes many more tokens.
Cite this pagev1.3.1 · 2026-09
Your number
What does your AI actually use?
Pick your model, set your usage, get the number, with sources.
Sources
Cited on this page, in order of appearance.
- 1Jegham, N., Abdelatti, M., Elmoubarki, L., & Hendawi, A. (2025). How Hungry is AI? Benchmarking Energy, Water, and Carbon Footprint of LLM Inference. arXiv preprint arXiv:2505.09598 (v6, Nov 2025). (accessed 2026-06) preprint
- 2Siddik, M. A. B., Shehabi, A., & Marston, L. (2021). The environmental footprint of data centers in the United States. Environmental Research Letters 16(6): watershed-scale direct + indirect water footprint. (accessed 2026-06) peer-reviewed
- 3You, J. (Epoch AI) (2025). How much energy does ChatGPT use?. Epoch AI Gradient Updates: GPT-4o per-query energy estimate ~0.3 Wh. (accessed 2026-06) industry report
- 4Altman, S. (2025). The Gentle Singularity. blog.samaltman.com: 0.34 Wh and 0.000085 gal of water per average ChatGPT query. (accessed 2026-06) official disclosure
- 5Google (2025). Measuring the environmental impact of AI inference. Google Cloud technical disclosure: 0.24 Wh / 0.26 mL / 0.03 gCO₂e per median Gemini Apps prompt. (accessed 2026-06) official disclosure
- 6Mistral AI (2025). Our contribution to a global environmental standard for AI. Mistral AI lifecycle analysis with Carbone 4 and ADEME, reviewed by Resilio and Hubblo: 45 mL water & 1.14 gCO₂e per 400-token Le Chat response (marginal inference); Mistral Large 2 training plus its first 18 months of use: 20.4 ktCO₂e and 281,000 m³ of water. No energy (Wh) figure disclosed. (accessed 2026-09) official disclosure
- 7AIWaterUse (2026). AIWaterUse methodology: disclosed site assumptions and derivations. aiwateruse.org/methodology: blended tokens per exchange, input/output split, unit definitions and arithmetic derivations, reviewed quarterly. (accessed 2026-06) site assumption