
As hyperscale data centers continue to grow both in size and power, with widespread backlash erupting over their presence and environmental footprint, greener cooling is no longer an option – it has to become a prerequisite. There have been a number of strides made in this area, such as using reclaimed wastewater. Amazon is using reclaimed water to cool some of its data centers and expects to use reclaimed wastewater at over 120 data centers by 2030. Other hyperscalers like Google, Meta, and Microsoft have also announced various water projects as part of a pledge to become water positive in the coming years. Nvidia's new closed-loop liquid cooling design could address some of the environmental concerns that hang over the data center and AI industries.
Data centers have historically used some form of open-loop evaporative cooling, usually combining chillers and cooling towers to cool warm air by water evaporation. This method reduces electricity consumption, but the trade-off is extreme water consumption, as this method requires a constant supply of makeup water. According to a report published at arXiv, between 70% to 80% of the water used for evaporative cooling is lost in the process. Closed-loop liquid cooling is often more sustainable but has faced a variety of obstacles that have slowed its adoption in data centers, while newer technologies like direct-to-chip cooling and two-phase immersion cooling show even more promise. Microsoft has already been using direct-to-chip cooling and rack-level heat exchangers with its Maia 100 platform, and has also been working on its microfluidics approach to direct-to-chip cooling.
Nvidia's take on direct-to-chip cooling uses 100 percent liquid cooling, and touts the potential for zero water consumption and no fans. This is because Nvidia's new AI server platform is designed to run a lot hotter than you'd expect.
Designed to run hotter than a hot tub at 45°C
Nvidia's new liquid-cooling design is part of the company's DSX AI platform, a reference AI factory design that outlines infrastructure built on a suite of Nvidia technologies. Within the DSX platform, Nvidia's methodology for direct-to-chip cooling is part of its DSX MaxLPS, which itself is a handful of technologies designed around performance per watt and maximizing compute across various power budgets.
Nvidia's closed-loop, direct-to-chip cooling is designed around a coolant mixture of 75 percent water and 25 percent propylene glycol, with a coolant temperature of 45 °C, and the system has an inlet temperature of 45°C and an outlet temperature of 55°C. When the liquid exits the rack, it's piped to an outdoor passive radiator that discharges the heat load. By using a loop that runs on warm water, a data center can theoretically avoid relying on mechanical cooling, such as chillers or cooling towers. This also means that the temperature delta between the warm water and the ambient air outside isn't extreme enough for a passive radiator to discharge it, though Nvidia is mindful to point out that this can be geographically dependent. Arizona, for instance, where summer temperatures can regularly rise into the triple digits, chillers will still be required.
This design also incorporates monolithic cold plates, which allow the liquid to pass over not only the GPUs and CPUs, but also the NVLink switches, optic modules, and the power bus bars, which are rated for a continuous current of 5,000 amps. The closed loop is designed to be filled once, and Nvidia claims it will last for the lifetime of the facility – a bold claim to be sure, and one that may or may not ease the concerns over data centers and water supply interference.
The use of warm water cooling is a departure from traditional approaches that rely on cold water to maximize heat transfer. By operating at higher temperatures, Nvidia can reduce or eliminate the need for energy-intensive chillers, which in turn lowers the overall power usage effectiveness (PUE) of the data center. However, the effectiveness of passive radiators depends heavily on local climate conditions. In cooler regions, the warm coolant can reject heat directly to the air without additional mechanical assistance. In hotter climates, supplemental cooling may still be needed, but the overall water consumption is drastically reduced compared to evaporative towers.
Nvidia's new design is mandatory, and the water nobody is talking about
Nvidia's move to 100 percent liquid cooling is not only a matter of resource management, but also one of necessity. Nvidia's previous Grace Blackwell platform already used liquid cooling for the GPUs, but the CPUs and other components relied on air cooling. Nvidia's Vera Rubin architecture pushes the power envelope to a different order of magnitude – one Rubin GPU uses as much as 2,300 watts, marking a 64 percent increase in TDP over the older GB300 GPU. A Vera Rubin NVL72 rack, which is composed of 72 Rubin GPUs and 36 Vera CPUs, in addition to other rack hardware (switches, optic modules, etc.), exceeds 200 kW – well beyond what air can cool.
Data centers with cooling towers use roughly 2.6 million gallons of water per megawatt per year, so while Nvidia's design that claims almost no water shouldn't be taken at face value, it also shouldn't be ignored. But it's only part of the equation, as a data center's on-site water usage doesn't paint the whole picture – AI's water problem extends beyond the data center to the power plant. US-based fossil fuel power plants consume an estimated 2.7 billion gallons of water per day, according to the U.S. Geological Survey. The International Energy Agency (IEA) estimates that coal and natural gas power plants will supply 40% of the increased electricity demand from data centers through 2030.
That means there's a lot of water being used down the street from data centers that those data centers don't usually count in their water footprint. Many estimates that try to quantify how much water AI uses also don't typically take power generation into account. Nvidia's full liquid cooling may reduce water usage at the facility level, but there's still a lot of work to do outside the data center.
Beyond water, the thermal management of high-density racks presents engineering challenges. The monolithic cold plates used in Nvidia's design must efficiently extract heat from components that generate enormous thermal loads. The power bus bars carrying up to 5,000 amps require careful design to avoid hot spots and ensure even cooling. The coolant mixture of water and propylene glycol is chosen for its thermal properties and freeze protection, but it also requires careful maintenance to prevent corrosion and microbial growth over the facility's lifetime. Nvidia's claim of a fill-and-forget system is ambitious; any leakage or degradation could compromise the entire cooling loop.
The broader context of data center sustainability involves not only cooling but also the source of electricity. Many hyperscalers are investing in renewable energy to power their facilities, which reduces both carbon emissions and water consumption at power plants. However, the rapid growth of AI compute demand means that new data centers are being built faster than renewable capacity can be added. In regions where the grid relies on fossil fuels, the net environmental impact may still be negative even with efficient cooling. Nvidia's design is a step forward, but it is not a silver bullet.
Industry experts note that the shift to liquid cooling is inevitable as chip power densities continue to rise. The Vera Rubin architecture is just the beginning; future generations of GPUs and CPUs will likely push beyond 3,000 watts per chip. This will require even more innovative cooling solutions, such as two-phase immersion cooling or dielectric fluids that can handle extreme heat fluxes. Nvidia's closed-loop approach provides a scalable foundation that can be adapted to these future needs. By proving that warm water cooling works at scale, Nvidia may accelerate adoption across the industry.
Meanwhile, other companies are exploring alternative cooling methods. Google has experimented with machine learning to optimize data center cooling and uses a combination of air and liquid cooling in its TPU pods. Meta has invested in direct-to-chip cooling for its AI servers, and Microsoft is developing two-phase immersion tanks for its Azure cloud. Each approach has trade-offs in cost, complexity, and sustainability. Nvidia's design emphasizes simplicity and low water use, but it requires a dedicated plumbing infrastructure and careful site selection.
Another consideration is the heat rejection system. Passive radiators, while energy-efficient, require large surface areas and may not be feasible in dense urban data centers. Active dry coolers or cooling towers could be added, but they would reintroduce water or energy consumption. Nvidia's reference design assumes ample outdoor space for radiators, which may not be available in all locations. The company recommends a site-specific analysis to determine the optimal cooling configuration.
The data center industry is under increasing pressure from regulators and communities to reduce its environmental footprint. In regions facing water scarcity, any technology that cuts water consumption is valuable. But the demand for AI compute is growing so rapidly that absolute resource use may still rise even with more efficient designs. Nvidia's MaxLPS platform addresses the immediate challenge of cooling high-density racks, but the long-term solution must involve a combination of efficiency improvements, renewable energy, and possibly new forms of computing that are less resource-intensive.
In conclusion, Nvidia's new data center cooling design is a significant technological leap. By operating at higher temperatures and using closed-loop liquid cooling, the company has demonstrated a path to zero water consumption at the facility level. However, the broader water and energy footprints of AI remain tied to the power grid and the manufacturing of hardware. As data centers become more efficient, we must not overlook the hidden costs that occur upstream and downstream. The future of sustainable AI depends on holistic thinking that encompasses every aspect of the infrastructure.
Source:SlashGear News
