Methodology: How We Estimate AI's Water, Energy, and Carbon
Every figure on this site is one multiplication, done three ways: the energy a prompt uses, times a water (or carbon) factor. The three ways (Low, Mid, High) are not optimism and pessimism. They are accounting boundaries: cooling water only, plus power-plant water, or the full lifecycle including training and chip manufacturing. Every coefficient below carries a source.
How an estimate is built
When you ask the calculator about a prompt, five things happen, in order:
- Your usage becomes tokens. A typical chat exchange (your prompt plus the model's reply) averages about 700 tokens (site assumption, consistent with the 400-token-response benchmarks in Jegham et al. 2025 plus a typical prompt). If you enter token counts directly, we use those.
- Tokens become energy. Each model is assigned to a class (efficient, standard, frontier, or reasoning) with a benchmarked energy intensity in watt-hours per 1,000 tokens (Jegham et al. 2025; operator disclosures where they exist, like Google's 0.24 Wh per median Gemini prompt).
- Energy becomes water. We multiply by liters per kilowatt-hour. This is where the scenario toggle matters: the factor ranges from 0.3 L/kWh (Low) to 30 L/kWh (High), and the next section explains why.
- Energy becomes carbon. Same multiplication, different factor: grams of CO₂e per kilowatt-hour, from clean-power operator (Low) to fossil-heavy grid (High), anchored to EPA eGRID.
- The result becomes something you can picture. Milliliters are abstract; bottles, showers, and cups of coffee are not. Every equivalent is sourced too.
The three scenarios, honestly
Ask "how much water does a chatbot query use?" and you can get answers from a quarter of a milliliter to half a liter (a factor of a thousand or so), all from credible sources. The numbers disagree because the questions disagree: each study draws its boundary around a different amount of the system. Rather than crown one boundary correct, we compute all three and label them. Think of it like asking what a hamburger costs: the menu price, the price with tax and tip, or the price including the farm subsidies. Same hamburger, three defensible answers.
Low: cooling only
What it counts: the water evaporated by the data center's own cooling systems, per unit of computing. Nothing else.
This is the boundary operators use, because it is the part of the system they own and meter. Google's official figure (0.26 mL per median Gemini prompt) lives here, as does OpenAI's ~0.32 mL per ChatGPT query (Altman, June 2025). Efficient operators report on-site water use of about 0.3 liters per kilowatt-hour (Microsoft's fleet average is 0.30). Google's own per-prompt figure implies about 1.1 L/kWh and OpenAI's about 0.9, so our 0.3 L/kWh Low factor sits at the efficient end of the operator range, and the calculator's Low values for Gemini and ChatGPT come out below those companies' own figures.
Low is a real number, honestly measured. It is also the answer to the narrowest possible question. If you've seen a headline that AI barely uses any water, it was built on this boundary.
Mid: + electricity
What it counts: cooling water, plus the water consumed generating the electricity the data center runs on.
Power plants are thirsty. Thermoelectric plants evaporate cooling water; hydroelectric reservoirs lose water to evaporation. On the average US grid, generating one kilowatt-hour consumes roughly 4.5 liters of water somewhere upstream (Siddik et al. 2021), several times what the data center itself evaporates. Add the on-site share and Mid works out to 5 L/kWh, or about 2.8 mL for a typical chat exchange on a standard-class model.
We treat Mid as the most decision-useful boundary: it counts everything that happens because you ran the query, while you ran it, without reaching back into construction and manufacturing. It is the default view in the calculator.
High: full lifecycle
What it counts: everything in Mid, plus a per-query share of the water used to train the model and to manufacture the hardware it runs on.
Training a frontier model consumes water for months before the first user shows up, and fabricating chips is among the most water-intensive manufacturing on Earth. Lifecycle studies spread those upstream costs across the queries the model eventually serves. Mistral's lifecycle analysis, the first from a model developer and externally reviewed, found 45 mL of water and 1.14 gCO₂e per 400-token response on this boundary. Mistral published no energy figure; at this site's assumed ~1.5 Wh for such a response, that implies roughly 30 L/kWh. The widely quoted ~519 mL per 100-word email (Washington Post with UC Riverside, 2024) is often filed here, but it counts cooling plus electricity at a high 2024 GPT-4 energy estimate, not training or manufacturing.
High is the boundary you want if your question is "what does AI as a system cost in water?" rather than "what does my query add at the margin?"
One subtlety: withdrawal versus consumption
Throughout this site, "water use" means consumption (water evaporated or otherwise made unavailable to its source watershed), not withdrawal, water borrowed and returned (as once-through cooling does). Studies and headlines mix these freely, which accounts for some of the wilder numbers in circulation. Where a source reports withdrawal, we say so.
How the scenarios combine with energy
One detail for careful readers. Our energy benchmarks also come as a low/mid/high spread, but that spread means something different: it reflects variance across studies and models, not an accounting boundary. To avoid multiplying two kinds of uncertainty into one opaque number, displayed water and carbon figures are computed from the Mid energy intensity times the scenario's water or carbon factor. The scenario toggle changes what gets counted, not how pessimistic we feel about the energy literature.
The coefficients
Every number rendered anywhere on this site traces to the table below, which is generated directly from the same file the calculator engine reads, so this page cannot drift from the estimates. Each row carries its source.
Coefficients version 1.3.1 · updated 2026-09
| What it counts | Factor (L/kWh) | Source | |
|---|---|---|---|
| Low: cooling only | cooling only | 0.3 | Efficient-operator on-site cooling: Microsoft reports a fleet-average WUE of ~0.30 L/kWh (down from ~0.49 in 2021). Google's fleet runs higher: its Gemini disclosure (0.26 mL per 0.24 Wh) implies ~1.1 L/kWh, and Altman (2025) 0.32 mL at 0.34 Wh implies ~0.94 L/kWh, so both operators' own per-prompt figures sit above this factor. |
| Mid: + electricity | cooling + electricity generation | 5 | ~0.5 L/kWh on-site + ~4.5 L/kWh indirect water consumed generating US-grid electricity (Siddik et al. 2021, Environ. Res. Lett.; US grid water intensity). |
| High: full lifecycle | full lifecycle: + training amortization + hardware manufacturing | 30 | Site derivation: Mistral's LCA discloses 45 mL of water per 400-token response (training and hardware amortized in) but no energy figure; at this site's assumed ~1.5 Wh for such a response (inside the Jegham et al. standard band) that works out to ~30 L/kWh. |
Energy intensity by model class
Watt-hours per 1,000 blended tokens (or per image / 5-second clip for media models). The low/mid/high spread here reflects variance across benchmarks, a separate axis from the scenario toggle. Displayed water and carbon use the mid energy value.
| Class | Low | Mid | High | Source |
|---|---|---|---|---|
| efficientGPT nano tier, Claude Haiku, Gemini Flash, Llama small | 0.1 | 0.3 | 0.6 | Jegham et al. 2025 (v6 Nov 2025): GPT-4.1 nano ~0.52 Wh/1K tok; nano-class models <0.3 Wh per short query. |
| standardClaude Sonnet, GPT-4.1/4o mini, Mistral Large, Llama 70B+ | 0.4 | 0.8 | 1.5 | Jegham et al. 2025 (v6 Nov 2025): GPT-4.1 mini ~0.450 Wh per 400-token exchange (~1.13 Wh/1K tok); GPT-4o mini ~1.44 Wh/1K tok. Mistral's LCA (45 mL, 1.14 gCO2e per 400-token Mistral Large response) discloses no energy figure; the ~1.5 Wh this site pairs with that response is a site assumption inside this band. |
| frontierGPT-4o/4.1/5-class, Claude Opus, Gemini Pro | 1 | 2 | 3.5 | Jegham et al. 2025 (v6 Nov 2025): GPT-4.1 ~0.871 Wh per 400-token exchange (~2.2 Wh/1K tok); Altman 'The Gentle Singularity' (June 2025) 0.34 Wh/query reflects shorter median queries. |
| reasoningo3-class, DeepSeek-R1, extended thinking modes | 3 | 7 | 15 | Jegham et al. 2025 (v6 Nov 2025): DeepSeek-R1 ~19-29 Wh per long-form query; o3 revised down to ~1.18 Wh per short query in v6 (about 6x below v1). Band is a site judgment call for long hidden chain-of-thought workloads, where unbilled reasoning tokens multiply energy per visible token. |
| imageSDXL/Midjourney-class, one generated image | 1.5 | 3 | 6 | Luccioni, Jernite & Strubell, 'Power Hungry Processing' (FAccT 2024): mean ~2.9 Wh per generated image for SDXL-class models, measured range ~0.06-11.5 Wh. |
| videoGen-3/Sora-class, one ~5-second clip | 25 | 150 | 1,000 | MIT Technology Review (May 2025; measurements by S. Luccioni): current-quality 5s CogVideoX clip ~3.4 MJ (~940 Wh); older low-quality model ~30 Wh. Closed commercial models (Sora, Veo) publish no figures; band spans the published measurements with mid near their geometric mean. |
Model assignments
Named models map to a class, with model-specific overrides where an official figure exists. Models with no published energy data appear with no class and produce no estimate. When we don't know, we say so.
| Model | Class | Source |
|---|---|---|
| ChatGPT (standard) | frontier | Class anchor: Jegham et al. 2025 (v6) GPT-4.1 ~2.2 Wh/1K tok; Epoch AI GPT-4o ~0.3 Wh/query; Altman (2025) 0.34 Wh/query. |
| ChatGPT (mini) | standard | Jegham et al. 2025 (v6): GPT-4.1 mini ~0.450 Wh per 400-token exchange (~1.13 Wh/1K tok); GPT-4o mini ~1.44 Wh/1K tok; inside the standard band, not the efficient one. |
| ChatGPT (nano) | efficient | Jegham et al. 2025 (v6): GPT-4.1 nano ~0.52 Wh/1K tok, within the efficient band. |
| ChatGPT (thinking / o3) | reasoning | Jegham et al. 2025 reasoning-model benchmarks; v6 (Nov 2025) revised o3 down to ~1.18 Wh per short query; the band reflects long hidden chain-of-thought workloads. |
| Claude Opus | frontier | Frontier-class assignment by parameter/latency tier; no official disclosure. Jegham et al. 2025 frontier band. |
| Claude Sonnet | standard | Standard-class assignment by tier; Jegham et al. 2025 mid-size band. |
| Claude Haiku | efficient | Efficient-class assignment by tier; Jegham et al. 2025 small-model band. |
| Gemini (app) | standard | Google 2025 Gemini technical disclosure (0.24 Wh / 0.26 mL / 0.03 gCO2e per median prompt).Google 2025 Gemini technical disclosure: 0.24 Wh per median Gemini Apps prompt; converted at this site's 700 blended tokens/prompt assumption to ~0.34 Wh/1K tokens. |
| Gemini Pro | frontier | Frontier-class assignment by tier; Google's official figure covers the median app prompt, not Pro-tier API usage. |
| Gemini Flash | efficient | Efficient-class assignment by tier; consistent with Google 2025 median-prompt disclosure. |
| DeepSeek-R1 | reasoning | Jegham et al. 2025: DeepSeek-R1 among the highest measured per-query energy (long reasoning chains). |
| Mistral Large | standard | Mistral AI LCA: 45 mL water & 1.14 gCO2e per 400-token response (marginal inference); no energy figure disclosed. Energy comes from the standard-class band; ~1.5 Wh per such response is a site assumption. |
| Llama (70B+) | standard | Jegham et al. 2025 open-weights 70B-class band. |
| Llama (small) | efficient | Jegham et al. 2025 small open-weights band. |
| Grok | frontier | Frontier-class assignment by tier; no official disclosure from xAI. |
| Image generation | image | Luccioni et al. (FAccT 2024): mean ~2.9 Wh per generated image, SDXL/Midjourney-class. |
| Video generation (5s clip) | video | MIT Technology Review (May 2025), Luccioni measurements: ~940 Wh per current-quality 5s clip; no official disclosures from Sora/Veo-class providers. |
| Claude Fable 5 | no data yet | Released June 2026 as a new Anthropic tier above Opus. No independent energy benchmarks (Jegham et al. methodology or otherwise) and no official disclosure exist yet, so it carries no class and produces no estimate. Listed so the absence is deliberate, not an oversight. |
Carbon intensity (gCO₂e per kWh)
| What it counts | Factor | Source | |
|---|---|---|---|
| Low: cooling only | clean-PPA operator | 50 | Operators with 24/7 clean power purchase agreements; EPA eGRID cleanest regional mixes. |
| Mid: + electricity | approx. US grid average | 350 | A rounded value slightly below recent EPA eGRID US averages (eGRID2022: 823 lb CO2/MWh, about 373 g/kWh). |
| High: full lifecycle | fossil-heavy grid | 500 | EPA eGRID fossil-heavy regional mixes. |
Assumptions
- Tokens per exchange (everyday mode): 700Site assumption, disclosed on /methodology: a typical chat exchange (prompt + response) averages ~700 blended tokens, consistent with the 400-token response benchmarks in Jegham et al. 2025 plus typical prompt length.
- Default input/output blend: 30% / 70%Site assumption, disclosed on /methodology: default 30/70 input/output blend for token-mode estimates.
- Home-page ticker rate: 230 L/secondQuarterly-reviewed constant. Derivation: ~2.5B ChatGPT queries/day (OpenAI disclosed, 2025) scaled to ~5B total AI queries/day (site estimate) x ~0.8 Wh per query x Mid 5 L/kWh = ~4 mL/query = ~20M L/day = ~230 L/second. The ~0.8 Wh is a blended site assumption, not a per-1K-token coefficient: it sits between this file's values for a ~700-token exchange on a standard-class model (~0.56 Wh) and a frontier-class one such as ChatGPT (~1.4 Wh). Operators' own per-prompt disclosures are lower (0.24-0.34 Wh); at standard-class energy alone the rate would be ~160 L/s. Always labeled 'estimated'.
Equivalents library
| Item | Value | Source |
|---|---|---|
| water bottle | 0.5 L | Standard 500 mL bottle (unit definition). |
| almond's worth of water | 4 L | ~4 L (1.1 gal) of irrigation water per almond, California average. |
| toilet flush | 6 L | Toilet at the US federal maximum of 1.6 gal (~6 L) per flush; WaterSense-labeled models use 1.28 gal or less. |
| shower | 65 L | ~8-minute shower at ~2.1 gal/min, EPA WaterSense. |
| cup of coffee (grown and brewed) | 130 L | One cup of coffee including water to grow the beans (Water Footprint Network). |
| day of a US person's home water use | 310 L | US person's home water use per day: 82 gal (EPA WaterSense). |
| hamburger's worth of water | 2,500 L | One hamburger with a 150 g beef patty, including feed and livestock water (beef ~15,400 L/kg, Water Footprint Network global average); a strict quarter-pounder is closer to ~1,800 L. |
| pair of jeans | 7,500 L | One pair of cotton jeans, full production chain (Water Footprint Network). |
| Olympic pool | 2,500,000 L | Olympic swimming pool, 2.5 ML (World Aquatics/FINA dimensions). |
| hour of an LED bulb | 0.01 kWh | 10 W LED bulb for one hour (arithmetic from device rating). |
| phone charge | 0.012 kWh | Full smartphone charge, ~12 Wh battery (arithmetic from battery spec). |
| hour of streaming | 0.08 kWh | One hour of streaming video, device + network + data center (IEA/Carbon Trust). |
| kettle boil | 0.11 kWh | Boiling a full kettle (~1 L; arithmetic from specific heat). |
| EV mile | 0.3 kWh | One mile in an average EV (~0.3 kWh/mi, EPA combined). |
| day of a US household's electricity | 29 kWh | Average US household electricity per day (~10.5 MWh/yr, EIA). |
| mile in a gas car | 0.4 kg CO₂e | One mile in an average US gasoline car (~404 g CO2/mi, EPA). |
| tree-year of CO₂ absorption | 21 kg CO₂e | CO2 absorbed by one mature tree in a year: a commonly cited ~21 kg round figure; published estimates range roughly 10-40 kg by species, age and climate. An order-of-magnitude equivalent, not a measured value. |
| transatlantic flight | 700 kg CO₂e | Economy seat, one-way transatlantic: ~700 kg CO2e, a site midpoint between fuel-only calculators such as ICAO's (~350-400 kg) and calculators that add non-CO2 warming effects (~1,000+ kg). |
frontier class · 2 Wh per 1,000 tokens (mid benchmark)
The same prompt can be “five drops” or “a bottle of water” depending on what you count: just data-center cooling (Low), the water behind the electricity (Mid), or training and hardware too (High).
Read the methodology16.9gallons
= 63.9 liters
ChatGPT (standard), 25 prompts/day · per year, Mid scenario
Energy
12.8 kWh
Carbon
4.5 kg CO₂e
What we don't count, and other honest limitations
- Your device and the network. The phone in your hand and the fiber between you and the data center use energy too. It's small per query and excluded from all three scenarios.
- Location. A liter evaporated in Phoenix is not a liter evaporated in Oslo. Our factors are US-grid averages; local reality varies by an order of magnitude in both directions.
- Recycled and non-potable water. Some operators cool with reclaimed wastewater or seawater. Our factors don't distinguish water quality, a genuine limitation of the available data.
- Staleness. Models, hardware, and grids all change faster than the literature. The coefficient file is versioned and dated, and we update it when better numbers are published, not when the discourse does.
How to disagree with us
Productively, we hope. Every coefficient lives in a single versioned file with a source attached, rendered above without editorial smoothing. If you think a number is wrong, you can identify exactly which coefficient, see exactly which study it came from, and point us to a better one. That is the entire trick: we'd rather be corrected than persuasive.