Skip to content
AIWaterUse.org
ES
Menu
↑ ↓ to move · Enter to open · Esc to close
MethodologyCoefficients version 1.3.1 · updated 2026-09Estimates changelog

Methodology: How We Estimate AI's Water, Energy, and Carbon

Every figure on this site is one multiplication, done three ways: the energy a prompt uses, times a water (or carbon) factor. The three ways (Low, Mid, High) are not optimism and pessimism. They are accounting boundaries: cooling water only, plus power-plant water, or the full lifecycle including training and chip manufacturing. Every coefficient below carries a source.

How an estimate is built

When you ask the calculator about a prompt, five things happen, in order:

  1. Your usage becomes tokens. A typical chat exchange (your prompt plus the model's reply) averages about 700 tokens (site assumption, consistent with the 400-token-response benchmarks in Jegham et al. 2025 plus a typical prompt). If you enter token counts directly, we use those.
  2. Tokens become energy. Each model is assigned to a class (efficient, standard, frontier, or reasoning) with a benchmarked energy intensity in watt-hours per 1,000 tokens (Jegham et al. 2025; operator disclosures where they exist, like Google's 0.24 Wh per median Gemini prompt).
  3. Energy becomes water. We multiply by liters per kilowatt-hour. This is where the scenario toggle matters: the factor ranges from 0.3 L/kWh (Low) to 30 L/kWh (High), and the next section explains why.
  4. Energy becomes carbon. Same multiplication, different factor: grams of CO₂e per kilowatt-hour, from clean-power operator (Low) to fossil-heavy grid (High), anchored to EPA eGRID.
  5. The result becomes something you can picture. Milliliters are abstract; bottles, showers, and cups of coffee are not. Every equivalent is sourced too.

The three scenarios, honestly

Ask "how much water does a chatbot query use?" and you can get answers from a quarter of a milliliter to half a liter (a factor of a thousand or so), all from credible sources. The numbers disagree because the questions disagree: each study draws its boundary around a different amount of the system. Rather than crown one boundary correct, we compute all three and label them. Think of it like asking what a hamburger costs: the menu price, the price with tax and tip, or the price including the farm subsidies. Same hamburger, three defensible answers.

Low: cooling only

What it counts: the water evaporated by the data center's own cooling systems, per unit of computing. Nothing else.

This is the boundary operators use, because it is the part of the system they own and meter. Google's official figure (0.26 mL per median Gemini prompt) lives here, as does OpenAI's ~0.32 mL per ChatGPT query (Altman, June 2025). Efficient operators report on-site water use of about 0.3 liters per kilowatt-hour (Microsoft's fleet average is 0.30). Google's own per-prompt figure implies about 1.1 L/kWh and OpenAI's about 0.9, so our 0.3 L/kWh Low factor sits at the efficient end of the operator range, and the calculator's Low values for Gemini and ChatGPT come out below those companies' own figures.

Low is a real number, honestly measured. It is also the answer to the narrowest possible question. If you've seen a headline that AI barely uses any water, it was built on this boundary.

Mid: + electricity

What it counts: cooling water, plus the water consumed generating the electricity the data center runs on.

Power plants are thirsty. Thermoelectric plants evaporate cooling water; hydroelectric reservoirs lose water to evaporation. On the average US grid, generating one kilowatt-hour consumes roughly 4.5 liters of water somewhere upstream (Siddik et al. 2021), several times what the data center itself evaporates. Add the on-site share and Mid works out to 5 L/kWh, or about 2.8 mL for a typical chat exchange on a standard-class model.

We treat Mid as the most decision-useful boundary: it counts everything that happens because you ran the query, while you ran it, without reaching back into construction and manufacturing. It is the default view in the calculator.

High: full lifecycle

What it counts: everything in Mid, plus a per-query share of the water used to train the model and to manufacture the hardware it runs on.

Training a frontier model consumes water for months before the first user shows up, and fabricating chips is among the most water-intensive manufacturing on Earth. Lifecycle studies spread those upstream costs across the queries the model eventually serves. Mistral's lifecycle analysis, the first from a model developer and externally reviewed, found 45 mL of water and 1.14 gCO₂e per 400-token response on this boundary. Mistral published no energy figure; at this site's assumed ~1.5 Wh for such a response, that implies roughly 30 L/kWh. The widely quoted ~519 mL per 100-word email (Washington Post with UC Riverside, 2024) is often filed here, but it counts cooling plus electricity at a high 2024 GPT-4 energy estimate, not training or manufacturing.

High is the boundary you want if your question is "what does AI as a system cost in water?" rather than "what does my query add at the margin?"

One subtlety: withdrawal versus consumption

Throughout this site, "water use" means consumption (water evaporated or otherwise made unavailable to its source watershed), not withdrawal, water borrowed and returned (as once-through cooling does). Studies and headlines mix these freely, which accounts for some of the wilder numbers in circulation. Where a source reports withdrawal, we say so.

How the scenarios combine with energy

One detail for careful readers. Our energy benchmarks also come as a low/mid/high spread, but that spread means something different: it reflects variance across studies and models, not an accounting boundary. To avoid multiplying two kinds of uncertainty into one opaque number, displayed water and carbon figures are computed from the Mid energy intensity times the scenario's water or carbon factor. The scenario toggle changes what gets counted, not how pessimistic we feel about the energy literature.

The coefficients

Every number rendered anywhere on this site traces to the table below, which is generated directly from the same file the calculator engine reads, so this page cannot drift from the estimates. Each row carries its source.

Coefficients version 1.3.1 · updated 2026-09

What it countsFactor (L/kWh)Source
Low: cooling onlycooling only0.3Efficient-operator on-site cooling: Microsoft reports a fleet-average WUE of ~0.30 L/kWh (down from ~0.49 in 2021). Google's fleet runs higher: its Gemini disclosure (0.26 mL per 0.24 Wh) implies ~1.1 L/kWh, and Altman (2025) 0.32 mL at 0.34 Wh implies ~0.94 L/kWh, so both operators' own per-prompt figures sit above this factor.
Mid: + electricitycooling + electricity generation5~0.5 L/kWh on-site + ~4.5 L/kWh indirect water consumed generating US-grid electricity (Siddik et al. 2021, Environ. Res. Lett.; US grid water intensity).
High: full lifecyclefull lifecycle: + training amortization + hardware manufacturing30Site derivation: Mistral's LCA discloses 45 mL of water per 400-token response (training and hardware amortized in) but no energy figure; at this site's assumed ~1.5 Wh for such a response (inside the Jegham et al. standard band) that works out to ~30 L/kWh.

Energy intensity by model class

Watt-hours per 1,000 blended tokens (or per image / 5-second clip for media models). The low/mid/high spread here reflects variance across benchmarks, a separate axis from the scenario toggle. Displayed water and carbon use the mid energy value.

ClassLowMidHighSource
efficientGPT nano tier, Claude Haiku, Gemini Flash, Llama small0.10.30.6Jegham et al. 2025 (v6 Nov 2025): GPT-4.1 nano ~0.52 Wh/1K tok; nano-class models <0.3 Wh per short query.
standardClaude Sonnet, GPT-4.1/4o mini, Mistral Large, Llama 70B+0.40.81.5Jegham et al. 2025 (v6 Nov 2025): GPT-4.1 mini ~0.450 Wh per 400-token exchange (~1.13 Wh/1K tok); GPT-4o mini ~1.44 Wh/1K tok. Mistral's LCA (45 mL, 1.14 gCO2e per 400-token Mistral Large response) discloses no energy figure; the ~1.5 Wh this site pairs with that response is a site assumption inside this band.
frontierGPT-4o/4.1/5-class, Claude Opus, Gemini Pro123.5Jegham et al. 2025 (v6 Nov 2025): GPT-4.1 ~0.871 Wh per 400-token exchange (~2.2 Wh/1K tok); Altman 'The Gentle Singularity' (June 2025) 0.34 Wh/query reflects shorter median queries.
reasoningo3-class, DeepSeek-R1, extended thinking modes3715Jegham et al. 2025 (v6 Nov 2025): DeepSeek-R1 ~19-29 Wh per long-form query; o3 revised down to ~1.18 Wh per short query in v6 (about 6x below v1). Band is a site judgment call for long hidden chain-of-thought workloads, where unbilled reasoning tokens multiply energy per visible token.
imageSDXL/Midjourney-class, one generated image1.536Luccioni, Jernite & Strubell, 'Power Hungry Processing' (FAccT 2024): mean ~2.9 Wh per generated image for SDXL-class models, measured range ~0.06-11.5 Wh.
videoGen-3/Sora-class, one ~5-second clip251501,000MIT Technology Review (May 2025; measurements by S. Luccioni): current-quality 5s CogVideoX clip ~3.4 MJ (~940 Wh); older low-quality model ~30 Wh. Closed commercial models (Sora, Veo) publish no figures; band spans the published measurements with mid near their geometric mean.

Model assignments

Named models map to a class, with model-specific overrides where an official figure exists. Models with no published energy data appear with no class and produce no estimate. When we don't know, we say so.

ModelClassSource
ChatGPT (standard)frontierClass anchor: Jegham et al. 2025 (v6) GPT-4.1 ~2.2 Wh/1K tok; Epoch AI GPT-4o ~0.3 Wh/query; Altman (2025) 0.34 Wh/query.
ChatGPT (mini)standardJegham et al. 2025 (v6): GPT-4.1 mini ~0.450 Wh per 400-token exchange (~1.13 Wh/1K tok); GPT-4o mini ~1.44 Wh/1K tok; inside the standard band, not the efficient one.
ChatGPT (nano)efficientJegham et al. 2025 (v6): GPT-4.1 nano ~0.52 Wh/1K tok, within the efficient band.
ChatGPT (thinking / o3)reasoningJegham et al. 2025 reasoning-model benchmarks; v6 (Nov 2025) revised o3 down to ~1.18 Wh per short query; the band reflects long hidden chain-of-thought workloads.
Claude OpusfrontierFrontier-class assignment by parameter/latency tier; no official disclosure. Jegham et al. 2025 frontier band.
Claude SonnetstandardStandard-class assignment by tier; Jegham et al. 2025 mid-size band.
Claude HaikuefficientEfficient-class assignment by tier; Jegham et al. 2025 small-model band.
Gemini (app)standardGoogle 2025 Gemini technical disclosure (0.24 Wh / 0.26 mL / 0.03 gCO2e per median prompt).Google 2025 Gemini technical disclosure: 0.24 Wh per median Gemini Apps prompt; converted at this site's 700 blended tokens/prompt assumption to ~0.34 Wh/1K tokens.
Gemini ProfrontierFrontier-class assignment by tier; Google's official figure covers the median app prompt, not Pro-tier API usage.
Gemini FlashefficientEfficient-class assignment by tier; consistent with Google 2025 median-prompt disclosure.
DeepSeek-R1reasoningJegham et al. 2025: DeepSeek-R1 among the highest measured per-query energy (long reasoning chains).
Mistral LargestandardMistral AI LCA: 45 mL water & 1.14 gCO2e per 400-token response (marginal inference); no energy figure disclosed. Energy comes from the standard-class band; ~1.5 Wh per such response is a site assumption.
Llama (70B+)standardJegham et al. 2025 open-weights 70B-class band.
Llama (small)efficientJegham et al. 2025 small open-weights band.
GrokfrontierFrontier-class assignment by tier; no official disclosure from xAI.
Image generationimageLuccioni et al. (FAccT 2024): mean ~2.9 Wh per generated image, SDXL/Midjourney-class.
Video generation (5s clip)videoMIT Technology Review (May 2025), Luccioni measurements: ~940 Wh per current-quality 5s clip; no official disclosures from Sora/Veo-class providers.
Claude Fable 5no data yetReleased June 2026 as a new Anthropic tier above Opus. No independent energy benchmarks (Jegham et al. methodology or otherwise) and no official disclosure exist yet, so it carries no class and produces no estimate. Listed so the absence is deliberate, not an oversight.

Carbon intensity (gCO₂e per kWh)

What it countsFactorSource
Low: cooling onlyclean-PPA operator50Operators with 24/7 clean power purchase agreements; EPA eGRID cleanest regional mixes.
Mid: + electricityapprox. US grid average350A rounded value slightly below recent EPA eGRID US averages (eGRID2022: 823 lb CO2/MWh, about 373 g/kWh).
High: full lifecyclefossil-heavy grid500EPA eGRID fossil-heavy regional mixes.

Assumptions

  • Tokens per exchange (everyday mode): 700Site assumption, disclosed on /methodology: a typical chat exchange (prompt + response) averages ~700 blended tokens, consistent with the 400-token response benchmarks in Jegham et al. 2025 plus typical prompt length.
  • Default input/output blend: 30% / 70%Site assumption, disclosed on /methodology: default 30/70 input/output blend for token-mode estimates.
  • Home-page ticker rate: 230 L/secondQuarterly-reviewed constant. Derivation: ~2.5B ChatGPT queries/day (OpenAI disclosed, 2025) scaled to ~5B total AI queries/day (site estimate) x ~0.8 Wh per query x Mid 5 L/kWh = ~4 mL/query = ~20M L/day = ~230 L/second. The ~0.8 Wh is a blended site assumption, not a per-1K-token coefficient: it sits between this file's values for a ~700-token exchange on a standard-class model (~0.56 Wh) and a frontier-class one such as ChatGPT (~1.4 Wh). Operators' own per-prompt disclosures are lower (0.24-0.34 Wh); at standard-class energy alone the rate would be ~160 L/s. Always labeled 'estimated'.

Equivalents library

ItemValueSource
water bottle0.5 LStandard 500 mL bottle (unit definition).
almond's worth of water4 L~4 L (1.1 gal) of irrigation water per almond, California average.
toilet flush6 LToilet at the US federal maximum of 1.6 gal (~6 L) per flush; WaterSense-labeled models use 1.28 gal or less.
shower65 L~8-minute shower at ~2.1 gal/min, EPA WaterSense.
cup of coffee (grown and brewed)130 LOne cup of coffee including water to grow the beans (Water Footprint Network).
day of a US person's home water use310 LUS person's home water use per day: 82 gal (EPA WaterSense).
hamburger's worth of water2,500 LOne hamburger with a 150 g beef patty, including feed and livestock water (beef ~15,400 L/kg, Water Footprint Network global average); a strict quarter-pounder is closer to ~1,800 L.
pair of jeans7,500 LOne pair of cotton jeans, full production chain (Water Footprint Network).
Olympic pool2,500,000 LOlympic swimming pool, 2.5 ML (World Aquatics/FINA dimensions).
hour of an LED bulb0.01 kWh10 W LED bulb for one hour (arithmetic from device rating).
phone charge0.012 kWhFull smartphone charge, ~12 Wh battery (arithmetic from battery spec).
hour of streaming0.08 kWhOne hour of streaming video, device + network + data center (IEA/Carbon Trust).
kettle boil0.11 kWhBoiling a full kettle (~1 L; arithmetic from specific heat).
EV mile0.3 kWhOne mile in an average EV (~0.3 kWh/mi, EPA combined).
day of a US household's electricity29 kWhAverage US household electricity per day (~10.5 MWh/yr, EIA).
mile in a gas car0.4 kg CO₂eOne mile in an average US gasoline car (~404 g CO2/mi, EPA).
tree-year of CO₂ absorption21 kg CO₂eCO2 absorbed by one mature tree in a year: a commonly cited ~21 kg round figure; published estimates range roughly 10-40 kg by species, age and climate. An order-of-magnitude equivalent, not a measured value.
transatlantic flight700 kg CO₂eEconomy seat, one-way transatlantic: ~700 kg CO2e, a site midpoint between fuel-only calculators such as ICAO's (~350-400 kg) and calculators that add non-CO2 warming effects (~1,000+ kg).
Try it with your own numbers
Model

frontier class · 2 Wh per 1,000 tokens (mid benchmark)

How much do you use it?25 prompts/day
Show results
What gets counted

The same prompt can be “five drops” or “a bottle of water” depending on what you count: just data-center cooling (Low), the water behind the electricity (Mid), or training and hardware too (High).

Read the methodology
WaterMid: + electricity

16.9gallons

= 63.9 liters

ChatGPT (standard), 25 prompts/day · per year, Mid scenario

Energy

12.8 kWh

Carbon

4.5 kg CO₂e

What we don't count, and other honest limitations

  • Your device and the network. The phone in your hand and the fiber between you and the data center use energy too. It's small per query and excluded from all three scenarios.
  • Location. A liter evaporated in Phoenix is not a liter evaporated in Oslo. Our factors are US-grid averages; local reality varies by an order of magnitude in both directions.
  • Recycled and non-potable water. Some operators cool with reclaimed wastewater or seawater. Our factors don't distinguish water quality, a genuine limitation of the available data.
  • Staleness. Models, hardware, and grids all change faster than the literature. The coefficient file is versioned and dated, and we update it when better numbers are published, not when the discourse does.

How to disagree with us

Productively, we hope. Every coefficient lives in a single versioned file with a source attached, rendered above without editorial smoothing. If you think a number is wrong, you can identify exactly which coefficient, see exactly which study it came from, and point us to a better one. That is the entire trick: we'd rather be corrected than persuasive.

Source bibliography

Every number rendered anywhere on this site traces to one of these entries.

official disclosure8

peer-reviewed4

government8

news7

Cite this pagev1.3.1 · 2026-09

Your number

What does your AI actually use?

Pick your model, set your usage, get the number, with sources.

Try the calculator

Sources

Cited on this page, in order of appearance.

  1. 1AIWaterUse (2026). AIWaterUse methodology: disclosed site assumptions and derivations. aiwateruse.org/methodology: blended tokens per exchange, input/output split, unit definitions and arithmetic derivations, reviewed quarterly. (accessed 2026-06) site assumption
  2. 2Jegham, N., Abdelatti, M., Elmoubarki, L., & Hendawi, A. (2025). How Hungry is AI? Benchmarking Energy, Water, and Carbon Footprint of LLM Inference. arXiv preprint arXiv:2505.09598 (v6, Nov 2025). (accessed 2026-06) preprint
  3. 3Google (2025). Measuring the environmental impact of AI inference. Google Cloud technical disclosure: 0.24 Wh / 0.26 mL / 0.03 gCO₂e per median Gemini Apps prompt. (accessed 2026-06) official disclosure
  4. 4Microsoft (2024). Sustainable by design: Next-generation datacenters consume zero water for cooling. Microsoft Cloud Blog, December 2024: fleet-average water usage effectiveness 0.30 L/kWh in the last fiscal year, down 39% from 0.49 L/kWh in 2021. (accessed 2026-09) official disclosure
  5. 5Mistral AI (2025). Our contribution to a global environmental standard for AI. Mistral AI lifecycle analysis with Carbone 4 and ADEME, reviewed by Resilio and Hubblo: 45 mL water & 1.14 gCO₂e per 400-token Le Chat response (marginal inference); Mistral Large 2 training plus its first 18 months of use: 20.4 ktCO₂e and 281,000 m³ of water. No energy (Wh) figure disclosed. (accessed 2026-09) official disclosure
  6. 6US Environmental Protection Agency (2024). Emissions & Generation Resource Integrated Database (eGRID). US EPA: grid carbon intensity by region. (accessed 2026-06) government
  7. 7Altman, S. (2025). The Gentle Singularity. blog.samaltman.com: 0.34 Wh and 0.000085 gal of water per average ChatGPT query. (accessed 2026-06) official disclosure
  8. 8Siddik, M. A. B., Shehabi, A., & Marston, L. (2021). The environmental footprint of data centers in the United States. Environmental Research Letters 16(6): watershed-scale direct + indirect water footprint. (accessed 2026-06) peer-reviewed
  9. 9Verma, P., & Tan, S. (The Washington Post), with UC Riverside researchers (2024). A bottle of water per email: the hidden environmental costs of using AI chatbots. The Washington Post, September 2024: ~519 mL of water and ~0.14 kWh per 100-word GPT-4 email (US average), counting data-center cooling plus the water behind electricity generation. (accessed 2026-09) news
  10. 10Luccioni, S., Jernite, Y., & Strubell, E. (2024). Power Hungry Processing: Watts Driving the Cost of AI Deployment?. FAccT 2024 (arXiv:2311.16863): mean ~2.9 Wh per generated image for SDXL-class models, measured range ~0.06-11.5 Wh. (accessed 2026-06) peer-reviewed
  11. 11O’Donnell, J., & Crownhart, C. (MIT Technology Review; measurements by S. Luccioni) (2025). We did the math on AI's energy footprint. MIT Technology Review, May 2025: ~3.4 MJ (~940 Wh) per current-quality 5-second CogVideoX clip; ~30 Wh for an older low-quality model. (accessed 2026-06) news
  12. 12You, J. (Epoch AI) (2025). How much energy does ChatGPT use?. Epoch AI Gradient Updates: GPT-4o per-query energy estimate ~0.3 Wh. (accessed 2026-06) industry report
  13. 13OpenAI (via TechCrunch) (2025). ChatGPT users send 2.5 billion prompts a day. TechCrunch, July 2025, reporting OpenAI figures: ~2.5B ChatGPT prompts per day; basis of the site’s scaled ~5B AI queries/day ticker estimate. (accessed 2026-09) official disclosure
  14. 14Water Footprint Network (Mekonnen, M. M., & Hoekstra, A. Y.) (2011). Product water footprint database. Water Footprint Network: agricultural water footprints (coffee, beef, cotton, almonds). (accessed 2026-06) peer-reviewed
  15. 15US Environmental Protection Agency (2024). WaterSense: residential water use and fixture flow rates. US EPA: 82 gal/person/day household use; fixture flow rates. (accessed 2026-06) government
  16. 16World Aquatics (FINA) (2023). Facility rules: Olympic swimming pool dimensions. World Aquatics facility rules: 50 m × 25 m × ≥2 m ≈ 2,500 m³ (2.5 ML). (accessed 2026-06) industry report
  17. 17Kamiya, G. (2020). The carbon footprint of streaming video: fact-checking the headlines. International Energy Agency commentary: ~0.08 kWh per viewing hour across device, network and data center. (accessed 2026-06) government
  18. 18US EPA / US Department of Energy (2025). fueleconomy.gov: electric vehicle energy consumption. EPA combined ratings: average EV ~0.3 kWh per mile. (accessed 2026-06) government
  19. 19US Energy Information Administration (2024). Electricity use in homes. US EIA: average US household electricity consumption ~29 kWh/day. (accessed 2026-06) government
  20. 20US Environmental Protection Agency (2024). Greenhouse gas emissions from a typical passenger vehicle. US EPA: ~404 g CO₂ per mile for an average US gasoline car. (accessed 2026-06) government
  21. 21AIWaterUse (2026). Annual CO₂ uptake of one mature tree (commonly cited estimate). aiwateruse.org/methodology: ~21 kg CO₂ per mature tree per year, a widely repeated round figure; published species- and site-specific estimates range roughly 10-40 kg/yr. An order-of-magnitude equivalent only. (accessed 2026-09) site assumption
  22. 22International Civil Aviation Organization (2024). ICAO Carbon Emissions Calculator. ICAO methodology (fuel-burn CO₂ only): roughly 350-400 kg CO₂ per economy seat, one-way New York-London. The site’s ~700 kg equivalent is a midpoint between fuel-only calculators and those adding non-CO₂ warming effects (~1,000+ kg). (accessed 2026-09) government