The hidden cost of cooling AI

The hidden cost of cooling AI

The Datacenter cooling debate has traditionally been framed as a simple engineering trade-off: use water, or use more electricity. AI is making that equation much harder, and much more expensive.

The latest generation of AI infrastructure is pushing operators towards direct liquid cooling, higher coolant temperatures, dry cooling, hybrid systems and increasingly sophisticated heat-rejection architectures. On the surface, the objective is obvious: remove more heat from increasingly dense racks while consuming less water, but executives should be wary of treating “low-water” or “zero-water” cooling as the end of the story; the hidden cost is that cooling decisions are increasingly becoming decisions about power capacity, grid infrastructure, capital expenditure, chemicals, resilience and even the wider water footprint of electricity generation.

This was made particularly clear when Schneider Electric published modelling on comparing four hypothetical 100 MW data-centre designs in Paris and Dallas which said moving to liquid cooling and operating with coolant supplied at 45°C could reduce water consumption by at least 50% in its Dallas scenario. It also argued that cooling towers can consume five to 20 times more water than dry coolers for an equivalent facility. The direction is compelling, but the economics are not one-dimensional.

Higher coolant temperatures can allow operators to reject heat to the atmosphere for more of the year, reducing reliance on mechanical refrigeration and evaporative systems – in fact NVIDIA has been pursuing the same principle: it’s latest AI infrastructure is designed for liquid entering the servers at temperatures up to 45°C. Whilst the result can be a dramatic reduction in operational water consumption, a dry or predominantly dry heat-rejection system is not free. Fans, pumps and mechanical equipment consume electricity – during periods of extreme heat, the facility may also face a choice between accepting higher power consumption, using supplemental cooling or allowing operating conditions to constrain performance.

As UC Riverside Professor Shaolei Ren put it recently, “there’s a pretty direct trade-off between how much water is used and how much energy is used” for temperature control. That trade-off should change how AI data-centre developers think about the cooling budget – water should no longer be viewed simply as a utility line item, neither should cooling electricity be buried inside PUE.

The real question is what combination of water, electricity, equipment, land and infrastructure produces the required thermal performance at the lowest total system cost?

Microsoft’s disclosures illustrate the progress, and the complexity – they reported average fleet-wide WUE of 0.27 litres per kilowatt-hour in 2025, down from 2.3 L/kWh in it’s earliest facilities. It also says approximately 90% of its owned fleet now operates with low, or zero-water cooling systems, however, eliminating water at the facility does not mean water is eliminated from the infrastructure chain.

Recent analysis from TechTarget highlighted what it called the “shadow water problem”: electricity generation itself can consume or withdraw substantial quantities of water. The analysis cited Lawrence Berkeley National Laboratory estimates that U.S. datacenters consumed about 176 TWh of electricity in 2023, with electricity generation associated with nearly 800 billion litres of water consumption, which is particularly important for AI campuses because their cooling architecture and power architecture are becoming inseparable.

A site designed around dry cooling may need greater electrical capacity precisely when ambient temperatures are highest, a site using evaporative assistance may reduce that electrical requirement but increase dependence on water availability – a site combining liquid cooling, batteries, renewable generation and flexible computing may be able to manage both, but at the cost of greater system complexity and capital investment. The industry is now responding accordingly.

NVIDIA recently launched its DSX Ready qualification programme, initially covering battery energy-storage systems and cooling distribution units, they explicitly framed power, cooling, water, site and grid constraints as interconnected considerations for AI factories. This is perhaps the most important development for executives: cooling is moving upstream in the infrastructure decision, it is no longer something to optimise after the electrical architecture and building have been designed.

There is also another hidden cost emerging: the materials required by advanced cooling systems.

A recent investigation reported growing demand for PFAS associated with semiconductor manufacturing and some advanced data-centre cooling technologies, this issue is particularly relevant to immersion systems using fluorinated fluids. The regulatory and liability implications are difficult to price today but the lesson is straightforward: replacing water consumption with another environmental or supply-chain exposure does not make the underlying constraint disappear. 

For executives planning the next multi MW AI campus, the strategic question is not whether to choose liquid cooling, it is whether the entire thermal system – chip, coolant, CDU, heat rejection, electrical supply, grid connection, water infrastructure and backup strategy has been designed as one system. The cheapest cooling system on a spreadsheet may not be the cheapest once the cost of grid reinforcement, peak electricity, water security, equipment redundancy and future regulation is included. 

AI infrastructure is quickly becoming a contest in thermal engineering, and increasingly, a contest in systems engineering. The winners will not necessarily be the operators who use the least water – they will be the operators who understand what that water saving actually costs, and where that cost has moved.