The Hidden Heat Tax Behind the AI Data Center Boom

The Hidden Heat Tax Behind the AI Data Center Boom

Every faster AI chip leaves more heat behind. The next infrastructure challenge is not simply finding electricity—it is moving that heat safely and efficiently, hour after hour.

Artificial intelligence may live in the cloud, but the machines behind it obey an unavoidable law of physics: almost every unit of electricity entering a server eventually becomes heat.

That heat has always existed in data centers. What is changing is its concentration. Modern AI systems pack dozens of powerful GPUs into tightly connected racks and run them for long periods. The result is not merely a larger electricity bill. It is an intense thermal load that older server rooms were never designed to handle.

This is why AI data center cooling is moving from the background of the technology industry to the center of infrastructure planning.

The AI Boom Comes With a Heat Bill

Data center electricity consumption is rising rapidly. In its 2026 update, the International Energy Agency reported that global data center electricity use increased 17% in 2025, while consumption at AI-focused facilities grew 50%. The agency projects that total data center electricity consumption could rise from about 485 terawatt-hours in 2025 to approximately 950 terawatt-hours in 2030.

The U.S. outlook is equally striking. Lawrence Berkeley National Laboratory estimated in June 2026 that data centers could account for 11.8% of U.S. electricity use in 2030, with possible outcomes ranging from 9.5% to 15.3%.

Those numbers are usually presented as a power-supply story. They are also a heat-removal story. A facility cannot keep adding high-performance processors unless its cooling system can carry away the heat those processors produce.

NVIDIA’s GB200 NVL72 illustrates the change in scale. The system combines 72 Blackwell GPUs and 36 Grace CPUs in a single liquid-cooled rack. NVIDIA specifies about 120 kilowatts of required cooling capacity for that rack. At this density, cooling is no longer something that can be solved by installing a few more fans or lowering the room temperature.

Cooling Is a Chain, Not a Single Machine

It is tempting to imagine cooling as an oversized air conditioner. A modern AI facility is more complicated.

Heat must first leave the chip. It then moves through a cold plate or another cooling interface, into a coolant loop, through pumps and a coolant distribution unit, and finally into the building’s larger heat-rejection system. Chillers, cooling towers, dry coolers, or a combination of technologies must then release that heat outside—or send it somewhere useful.

The entire chain must operate continuously. If one stage is undersized, the GPUs may slow themselves to prevent damage. If cooling stops, workloads may have to shut down even when electricity remains available.

This creates a practical distinction that is often missed: a property can have enough grid capacity on paper but still be unsuitable for the planned computing equipment. The building may lack the piping, water flow, mechanical space, structural support, controls, or redundancy required for high-density liquid cooling.

Three Cooling Approaches—and Why They Are Not the Same

The transition away from conventional air cooling is not one single technological shift. Data center operators are choosing among several designs.

Air cooling

Fans move air across server components, while computer-room cooling equipment removes the heat from the building. This remains effective for many ordinary cloud and business workloads. Its weakness appears when too much heat is concentrated in too little space: fans consume more electricity, hot spots become harder to control, and more room is needed to deliver sufficient airflow.

Direct-to-chip liquid cooling

A metal cold plate sits directly on a GPU or CPU. Water or another coolant circulates through sealed channels in the plate and carries heat away without touching the electronics. This approach can support much denser hardware than air cooling and is increasingly common in new AI systems.

The water in this system is part of a controlled, closed loop. It should not be confused with placing an entire server in water.

Immersion cooling

In immersion cooling, the server or its components are placed in an electrically nonconductive dielectric fluid. The fluid absorbs heat directly from the equipment and transfers it to a heat exchanger. The secondary side of that exchanger may then use a separate water loop.

Although perfectly pure water has very low electrical conductivity, it quickly absorbs ions and contaminants when exposed to air, metals, and circuit boards. For that reason, mainstream immersion systems use stable dielectric fluids rather than ordinary or purified water in direct contact with the server.

Immersion cooling can support extremely high computing density while reducing fans, air-handling equipment, cooling energy, and on-site water consumption. It is an advanced and promising technology. Some installations still require specialized equipment, compatible components, and different servicing procedures; environmental regulation is mainly a concern for certain fluorinated fluids used in some two-phase systems, not for every form of immersion cooling.

The Water Question Is More Complicated Than It Looks

“Liquid cooling” does not automatically mean heavy water consumption.

In an evaporative cooling tower, some water is intentionally evaporated to remove heat. In a sealed direct-to-chip or immersion loop, the same working fluid can circulate repeatedly. The facility can then use dry coolers or other heat-rejection equipment that consumes little or no water during normal operation.

Microsoft said in December 2024 that its newer chip-level cooling design uses a closed loop and consumes zero water for cooling during operation. The company estimated that this approach could avoid more than 125 million liters of water annually at each data center compared with an evaporative design.

However, low on-site water use does not mean zero overall water impact. Electricity generation and equipment manufacturing may also consume water. Dry cooling can save water but may require more electricity or larger equipment in hot weather. Evaporative cooling can reduce electrical demand but use more water.

Location therefore matters. A design that makes sense in a cool, water-rich region may be a poor choice in a hot, drought-prone area. Berkeley Lab researchers found in 2025 that water consumed per computing workload can vary by more than 10,000 times because of differences in server efficiency, utilization, local electricity generation, climate, and cooling technology.

Efficiency Should Be Measured by Computing Delivered

The most efficient-looking building is not necessarily the most efficient AI system.

Operators often track Power Usage Effectiveness, or PUE, which compares a facility’s total electricity use with the electricity used by its computing equipment. Water Usage Effectiveness, or WUE, measures water consumed relative to IT energy.

Both are useful, but neither tells the entire story. An efficient cooling system attached to poorly utilized servers can still waste resources. A better assessment considers how much useful computing work the facility delivers for each unit of electricity and water consumed.

This is especially important for AI because workload scheduling can change demand. Some training jobs may be shifted to cooler hours, regions with cleaner electricity, or periods when the grid has more capacity. Thermal storage and intelligent controls can also reduce peak cooling demand without reducing total computing output.

Waste Heat Could Become an Asset

Most data centers spend money to move heat outdoors and then discard it. Higher-temperature liquid cooling creates an opportunity to use some of that heat instead.

Warm water leaving an AI facility may help heat nearby offices, apartments, greenhouses, or industrial processes. This will not work everywhere: the customer must be close enough, the heat temperature must be useful, and demand should align with data center operations. But when the conditions are right, heat reuse can reduce both cooling demand and another building’s fuel consumption.

The idea changes the basic question from “How do we get rid of the heat?” to “Who nearby can use it?”

What Will Separate Successful AI Data Centers

The strongest projects will not treat chips, electricity, cooling, water, and location as separate decisions. They will design them as one system.

That means choosing hardware density that the building can support, matching the cooling method to local climate and water conditions, providing backup capacity, monitoring both energy and water, and planning how equipment will be maintained before the first rack is installed.

Electricity remains a major constraint on AI growth. But a power contract alone does not create a functioning AI data center. The facility must also be capable of collecting, transporting, and releasing an extraordinary amount of heat without interruption.

The hidden cost of the AI boom is not only the electricity consumed. It is the infrastructure required to handle everything that electricity leaves behind.

Sources and Further Reading