I've been hearing that a single ChatGPT message uses a full bottle of water for about a year now, from every corner of the internet you'd care to name. Not "some water, somewhere, eventually, once you trace the whole electricity chain back to a power plant" – a full bottle, per message, like the thing's got a drip tray fitted. Ask anyone where they got it from and the answer's always some limp version of "everyone knows that," which is the exact phrase people reach for when they've never once had to defend the number and have no intention of starting now, because defending it would mean admitting they picked it up from a fifteen-second clip and never checked again.
I'm not writing this to have a go at any one person who's said it to me, because there are too many of them now to single anyone out. I'm self-taught, five years of it, and a decent chunk of what I ship day to day runs through the same infrastructure everyone's currently frothing about – I write my TypeScript and Python and Svelte by hand, but for the languages I'm shakier in, C, Rust, Kotlin, I lean on Claude Code and a couple of other CLI tools to move faster, while I keep the architecture and protocol design to myself. I've also tried running models locally on my own hardware, out of the same curiosity that got me into self-hosting everything else, and the results have ranged from mediocre to actively disappointing – which if anything makes me less inclined to defend any of this out of misplaced loyalty, not more. So this isn't a hypothetical stake for me. I'm about to go looking for paid work built on exactly this stack, and I already know I'll spend a chunk of every job doing the same unpaid labour I do in every kitchen and comment section now: correcting a number that was wrong from the moment it was published, repeated by people who trusted whoever handed it to them, sitting downstream of a media ecosystem that has worked out it makes more money from the wrong number than the right one.
It's Not All AI in There
Before any of the actual numbers, there's a framing problem sitting underneath this entire argument, and it's a stupid one, because it's so easy to check.
Basically the entire internet runs in data centres. Not "some of it, if you squint" – all of it, structurally, as a matter of how the internet is built. Every bank transfer. Every government record. Every email. Your Netflix queue, your Spotify library, the photos backed up off your phone the moment you took them, the website you're reading this on right now, this blog itself, sitting on servers I don't own in a building I've never seen and never will. None of that is new. None of it needed a single line of AI code to exist, and most of it predates AI being a mainstream word by decades. JLL's own 2026 market outlook puts AI at only around a quarter of all data centre workloads in 2025, with the other three-quarters being the same enterprise storage, web hosting, and cloud compute that's been quietly holding the internet together since before most of the people angry about it were born. Of the roughly twelve thousand active data centres worldwide as of mid-2026, plenty aren't running a single AI workload. They're just doing what data centres have always done, which is keep the internet standing up while nobody thinks about them.
So when someone tells me they're furious about "AI data centres" while streaming a show, refreshing their banking app, and posting about it on a platform whose entire backend is the exact same category of building, I don't think they're lying. I think they've been handed a villain that's easier to point at than the actual scale of what they're already relying on every single hour of every day, and nobody's bothered correcting them, because "the internet you already depend on runs on this too" is a much worse headline than "AI is drinking the aquifer." The AI data centres didn't invent this industry. They just gave a decades-old, largely invisible piece of infrastructure a face people could finally get angry at, and the fury landed on the newest tenant instead of the building.
The Number That Started It
Go back to 2019. MIT Technology Review ran a piece by Karen Hao reporting that training a single AI model could emit as much carbon as five cars over their entire lifetimes. It was a genuinely shocking figure, and it tore through the discourse the way shocking figures do.
The research behind it came from a team at UMass Amherst, and the paper is one of the most cited pieces ever written on AI and the environment. It's also wrong, by a factor of roughly 90, for three specific and fairly mundane reasons, as a later paper correcting the numbers spelled out in detail nobody bothered reading. The Google training run they were estimating tested thousands of candidate models but only fully trained the promising ones – the UMass team assumed every single candidate got the full treatment. They assumed the training ran on ordinary graphics cards, when Google was using specialised TPU chips built for exactly this job. And they assumed an average commercial data centre's cooling overhead, around 60% losses, when Google's own centres were losing closer to 10%. Stack those three mistakes and a number that should have looked like two researchers' flight to a conference got dressed up as a fleet of cars, and nobody who mattered checked the working before it went everywhere.
Jeff Dean at Google has publicly asked the lead author to stop citing the figure, and has put the actual inflation factor at over 3,000x once you account for how much more efficient training has become since. The paper hasn't been updated. Neither has the 2021 "Stochastic Parrots" paper that leaned on the same number and now has over 13,000 citations, making it the single most-cited piece of AI ethics writing that exists. The original UMass study has roughly 2,700 citations. The paper that corrected it has about 1,600. That gap – a wrong number cited nearly twice as often as its own correction – isn't a coincidence, and it isn't really about AI. It's about which kind of finding a journalist, an academic, or someone with a ring light actually wants their name attached to. Credit where it's due: I owe most of this specific citation trail to Andy Masley's writing on the data centre panic, which is worth reading in full and did the archaeology I'm just repeating here.
Water Doesn't Work the Way the Headlines Imply
The water figures follow almost exactly the same pattern, and water specifically seems to catch people because it feels more visceral than electricity – a kilowatt-hour is abstract, a bottle of water is not, and there's something that clearly offends people about an ephemeral chatbot answer having a physical, wet cost attached to it. Offence is not the same thing as accuracy, and a lot of what's circulating trades entirely on the former.
Start with what "water use" at a data centre actually means, because the term is doing an enormous amount of unearned work in every headline that uses it. There are three separate things bundled under it, and almost nobody sharing the number bothers separating them. Most water pumped in for cooling gets returned straight to the municipal supply – that's non-consumptive, and it's the majority of the volume anyone's quoting when a headline says a facility "uses millions of gallons." Some of it evaporates during cooling and genuinely leaves the local water system – that's the one real, consumptive cost, and it varies wildly by cooling method and climate; a techUK survey of over seventy UK sites in 2025 found just over half already run waterless cooling entirely, and 64% of the rest use under 10,000 cubic metres a year, which is a long way from the framing you usually see. And then there's the water used generating the electricity in the first place, which is real but is a property of the grid mix, not of AI specifically – the same criticism applies to running a kettle, and nobody's writing furious posts about kettles.
Then there's the per-query figure, the one that actually gets shared and screenshotted and never checked. The commonly cited claim was that a single ChatGPT query uses about 3 watt-hours, roughly ten times a Google search. That number came from a 2023 estimate built on GPT-3.5 running on older hardware, assuming a wildly generous output length. Epoch AI redid the maths in 2025 with current hardware and realistic response lengths and landed on about 0.3 watt-hours – ten times lower, and roughly in line with a Google search rather than ten times worse than one. Sam Altman independently gave almost the identical figure, 0.34 Wh, comparing it to what an oven draws in a bit over a second. The 3 Wh number is still the one doing the rounds, because it made headlines first, and the correction – as with the emissions figure – barely made a sound on the way out.
I want to be precise about what I'm not saying here, because I'd rather not be lumped in with the people doing the exact same sloppy thing in the opposite direction. Newer reasoning models are genuinely hungrier – some estimates put a complex GPT-5 query closer to 18 to 40 Wh – and multiplied across billions of daily queries the totals are not trivial. I'm not arguing AI uses no resources. I'm arguing that the specific numbers circulating in the "everyone knows that" category are frequently a decade-old guess dressed up as current fact, repeated by people who would be furious if you did this to them about literally anything else they cared about.
I'm Not Claiming AI Is Resource-Benign
There's another distinction I want to make before anyone reads the preceding sections as an argument for "therefore AI is fine", because it isn't. I don't think large language models are universally benevolent in their resource usage. I think some of the numbers people use to criticise them are wrong, sometimes spectacularly wrong, and that those corrections matter. But a bad statistic doesn't turn into a good externality just because the statistic was bad.
There is a perfectly reasonable environmental argument against building ever-larger inference infrastructure, training ever-larger models, and running increasingly compute-heavy reasoning workloads for problems that sometimes didn't need an AI system in the first place. If billions of queries become tens or hundreds of times more computationally expensive because we're asking models to reason for longer, that's a resource question worth taking seriously. So is the hardware supply chain, the electricity required at scale, the physical infrastructure needed to house it, and the fact that demand for compute doesn't magically stop existing just because one particular viral water statistic turned out to be rubbish.
And, honestly, I'd much rather run this stuff locally if local models didn't currently run like dog shit for the things I actually need them to do on my reasonably powerful hardware.
That isn't a contradiction. It's basically the reason I've tried local models in the first place. I'm already inclined towards self-hosting, local infrastructure and keeping computation under my own control where the practical trade-offs make sense. If I could put a capable model on my own hardware, run it efficiently, get the quality and latency I need, and avoid sending every request halfway across the world to someone else's GPU cluster, I'd have a pretty obvious reason to do exactly that. The problem is that "just run a local LLM" is still a much better proposition in theory than it is for a lot of actual workloads. The models I can run locally on hardware I can reasonably own have, in my experience, ranged from useful to mediocre to actively frustrating, depending on the task. The cloud models are often materially better at the things I'm actually asking them to do.
So I'm not making a universal claim that centralised AI infrastructure is environmentally superior, or even that it should be. I'm making a narrower claim: if we're going to have that argument, we should at least use numbers that describe reality.
There is a particularly annoying tendency in this debate to treat those positions as mutually exclusive. If you point out that "a bottle of water per prompt" is a misleading way to describe water consumption, people hear "AI has no environmental cost". If you point out that newer reasoning models can consume dramatically more compute than simple queries, people hear "AI is destroying the planet". Both reactions are lazy. The interesting question is what the actual marginal resource cost is for a particular workload, on a particular system, using a particular cooling and electricity setup, compared with the alternative way of accomplishing the same thing.
Sometimes the answer will favour an LLM. Sometimes it absolutely won't.
And that matters because I'm not particularly interested in defending AI as a moral good. I use these systems because they're useful. I would prefer them to become more efficient, more local, more transparent about their resource consumption and less dependent on enormous centralised infrastructure where the workload doesn't justify it. If local models eventually become good enough to replace the cloud models I use, I'll be more than happy to move more of my workload onto my own hardware.
Until then, pretending that every cloud inference is environmentally catastrophic because someone put "one bottle of water" in a TikTok caption isn't serious environmental analysis. But pretending that correcting that caption means the resource question has disappeared isn't serious analysis either.
What's Actually True, Because Some of It Is
Here's where I refuse to let this collapse into "well actually it's all fine then," because it isn't, and the people doing genuine reporting on this deserve better than getting lumped in with the bottled-water crowd purely because they share a subject.
The grid strain is real. US data centres pulled around 183 terawatt-hours in 2024, over 4% of total national electricity use, and the Department of Energy's own projections put that share at 6.7 to 12% by 2028. That growth is outpacing the rest of demand by a wide margin, and it's landing hardest in specific places rather than being spread evenly across the country. In Manassas, Virginia, a resident called John Steinbach opened a January bill for $281, nearly triple what he'd paid the month before, and he's not alone in tracing that spike back to data centre load growth pushing utilities like Dominion Energy toward rate increases that fund grid upgrades mostly benefiting the commercial customers driving the demand in the first place. That's not a made-up grievance. That's someone's actual heating bill, and being angry about it is entirely correct.
Local water stress is also real in specific places, even where the "millions of gallons" framing usually swallows the non-consumptive majority whole. A February 2026 Water Research Centre estimate put England's data centre potable water consumption at close to 1.9 million cubic metres a year and climbing. Construction-phase issues – dewatering, sediment runoff during building work rather than ongoing operational use – have caused genuine local harm in places like Indiana, where state officials investigated whether groundwater pumped out to build a facility left nearby wells dry. None of that is imaginary, and none of it gets fixed by me pointing out that a chatbot query uses less water than the headline implies.
And the backlash has teeth for a reason. A Gallup poll from March 2026 found 70% of respondents opposed new data centre construction in their own neighbourhood, and local opposition blocked or delayed at least 75 US projects worth $130 billion in the first quarter of 2026 alone. Some of that opposition is running on inflated numbers. Not all of it is. Traffic, noise, light pollution, and genuine grid competition with residential users are ordinary planning concerns that would apply to any large industrial development, AI-branded or not, and treating every objector as a mark who fell for a viral stat is its own kind of contempt.
So the honest position isn't "the critics are all wrong" and it isn't "the industry is all lying." It's that a real, quantifiable set of local harms is currently buried under a pile of numbers that are wrong by an order of magnitude or more, and untangling the two requires reading past the headline, which is precisely the bit nobody involved – not the sharer, not the outlet, not the algorithm feeding them both – has any incentive to do.
Nobody Reads Past the Headline, and the Industry Knows It
This is the part that actually gets under my skin, more than any individual statistic.
Science journalism in the US lost about a quarter of its newsroom jobs between 2008 and 2020, and the specialist science desks that used to exist at most major papers have mostly been gutted into general assignment reporting, as this history of the panic covers in more depth. That's not a neutral fact about media economics I'm mentioning for colour. It means the reporter writing about data centre water use this month is quite possibly the same person who covered a council meeting last month, working against a deadline, inside a business model that pays out in shares regardless of whether the piece's own body actually supports its headline. Two of the most widely shared articles on data centre water use – one from the New York Times, one from the BBC – both cover incidents that turned out to be construction-phase runoff rather than operational water draw, a distinction the articles themselves make clear if you actually read to the end. Almost nobody sharing them does. Nobody's incentivised to.
That asymmetry – alarming claim spreads, correction dies quietly – isn't unique to AI, which is the bit that should actually worry you more than the AI part does. It happened to online video, to emails, to the "dig more coal, the PCs are coming" panic Forbes ran in 1999 about Amazon.com supposedly burning a lump of coal every time someone ordered a book. Every wave of new computing infrastructure gets the identical treatment, because "this ordinary-looking thing you use every day is secretly doing something enormous and hidden" is a genuinely satisfying thing to learn, the same way finding out a food you eat is secretly terrible for you is satisfying. It feels like being let in on something the people around you haven't clocked yet.
And I'm going to name the specific platform doing the most damage here, because pretending it's a vague, ambient "short-form video" problem lets the worst offender off far too easily: TikTok. A fifteen-second clip has room for the shocking claim and absolutely no room for the three caveats that would make it true, and that's before you even get to the comment section underneath it, which is where whatever's left of the nuance goes to be finished off entirely. I've spent enough time down there to know exactly what it looks like: nobody's citing a source, nobody's asking for one, and the top comment is never the correction, it's whichever unearned confident restatement of the original claim got the most thumbs fastest. Ask for a source in a TikTok comment section and you will, more often than not, get either silence, a screenshot of a different equally unsourced TikTok, or someone telling you to "just Google it" as though that isn't the exact instruction they themselves have never once followed. It's not that the users are stupid. It's that the entire format actively punishes anyone who stops to check anything, and rewards, with visibility and validation, whoever states the wrong number with the most conviction and the least hesitation. A platform that structurally cannot fit a caveat was never going to produce a well-calibrated public, and I don't think that's an accident so much as a business model.
And "Everyone" Might Not Even Be People Any More
Here's the bit that makes the whole cycle worse, and it's not a tangent, because it's the actual mechanism carrying the bad number around: a meaningful and rising share of "everyone" repeating "everyone knows that" isn't a person.
The exact figure depends entirely on who's counting and how, which I'm not going to smooth over just because it's inconvenient to my own argument. Cloudflare's own public radar put bots at roughly a third of all web traffic in mid-2026, humans still a clear majority. Imperva's and Thales's Bad Bot Reports, measuring differently, put automated traffic at 53% for 2025, officially overtaking humans. And in June 2026, Cloudflare separately reported that 57.4% of requests across a specific selection of the sites it hosts were automated, its own CEO admitting the crossover had happened faster than he'd predicted. Take the low end and it's still one in three requests. Take the high end and humans are, on that particular measure, no longer the majority of what's happening on the internet at all. Dead internet theorystarted as a half-joking forum conspiracy about most online content being bot-generated. It isn't a joke any more; the actual disagreement among people who measure this for a living is only ever about the exact percentage, never about the direction.
That matters here specifically, not just in the abstract. "Everyone knows that" was never a reliable source of truth even when everyone was definitely a person – it just meant a claim had spread, not that it was checked. Now some real, rising fraction of the accounts repeating it, liking it, and pushing it back up the feed toward the next person aren't checking it because they're not capable of checking anything. They're volume. A wrong number doesn't need to convince a single new human being to keep circulating; it just needs enough bot traffic amplifying it to look like consensus by the time a real person sees it for the third time and decides it must be true, because how could that many sources be wrong. The dead internet doesn't need to be entirely dead to break your sense of what "widely believed" actually means. It just needs to be dead enough that you can no longer tell the difference from inside it.
Why the Illiteracy Actually Costs Something
I don't think believing a wrong number about ChatGPT's water use is, on its own, some grave sin. What actually costs something is what happens once that wrong number becomes the load-bearing assumption a community, a council, or a national conversation builds its entire position on top of.
When 70% of people oppose the data centre near them, and a meaningful chunk of that opposition is running on a figure that's off by an order of magnitude, you don't get a well-calibrated planning process. You get projects blocked or waved through almost at random relative to their actual local impact, because the debate was never conducted on real numbers to begin with. Genuine grid strain and genuine rate hikes – the Steinbach bill, the Dominion Energy rate case – get buried under the exact same pile as the bottled-water claim, and the pile as a whole becomes trivially easy for an industry PR team to dismiss wholesale, because a good chunk of it genuinely is dismissible. That's the actual cost of technological illiteracy here, and it should make you angrier than the wrong number itself: it doesn't just embarrass the people who believed it, it degrades the entire public conversation's ability to tell a real problem from an invented one, at exactly the point where telling the two apart is the whole job.
And the industry is not innocent in any of this, so I'm not letting it off either. Plenty of hyperscaler messaging is just as happy to wave away genuine local harms by pointing at how overstated the viral numbers are, which is its own kind of bad faith dressed up as correction. "Your headline figure was wrong" is not the same claim as "there is no problem," and collapsing the two into one is exactly the move I'd expect from someone who'd rather the whole conversation just stopped, permanently, before anyone got round to the parts that are actually true.
I'm about to go and ask people to trust me with real technical work, for money, having spent five years teaching myself most of what I know outside a classroom, most of it maintaining the exact kind of infrastructure – protocol SDKs, self-hosted servers, a small personal data centre's worth of Nix-managed boxes on my own desk – that this entire panic claims to be about while mostly not being about at all. Part of that job, whether it's written into the offer or not, is going to be sitting across from someone who's absolutely certain about a number they've never once checked, and doing the unglamorous, thankless work of walking it back without making them feel stupid for having believed it, because they usually didn't get there by being stupid. They got there because the correction never had the marketing budget the original claim did, and it never will, and that's not a them problem. It's an information ecosystem built, deliberately, to reward being first and shocking over being right and boring, and I don't have a fix for that beyond writing another one of these and watching it lose to the bottled water anyway.