The coolant distribution unit is the boundary between the servers and the building. Most disputes on a liquid-cooling project are traceable to that boundary being described loosely in the purchase documents.
Fix one duty point before collecting quotes
Capacity means nothing without the temperatures it was measured at. State secondary supply and return, primary supply and return, and design flow, then require every bidder to quote net capacity at exactly those conditions with pump power included.
Describe redundancy as a drawing, not a letter
Ask each vendor to mark every pump, heat exchanger, valve, filter, controller, and power feed needed to carry design load with one component unavailable. A redundancy class in a datasheet rarely survives that exercise unchanged.
Water quality is a contract term
Secondary loop cleanliness determines cold-plate life, and responsibility for achieving it is frequently left ambiguous. Specify filtration, target particulate and chemistry limits, flushing procedure, acceptance testing, and who signs off.
- Filtration rating and maintenance interval
- Chemistry limits and permitted inhibitors
- Flushing and acceptance criteria before servers connect
- Ongoing sampling cadence and reporting
Controls and integration
Name the protocol, the points list, the alarm set, and the building system that owns setpoints. Establish what the unit does on loss of communication, loss of power, and a leak signal, and test each in commissioning.
What a CDU does
A coolant distribution unit separates the clean secondary loop that touches the servers from the facility water on the primary side, and it controls the flow rate, temperature, and pressure delivered to the cold plates. In doing so it takes on the thermal management duty that a data center previously handled entirely with air: the unit decides what temperature the silicon sees, and its controls are what keep that stable as the compute load swings.
In-rack, in-row, and multi-megawatt units
CDUs are sold in three broad formats. In-rack units sit inside the cabinet and suit pilots, single high-density racks, and halls without facility water at every row. In-row units serve several racks and are the common choice for a liquid-cooled pod. Large modular plants serve megawatts at a time and are where most AI data center capacity is now being specified. The formats differ in deployment granularity and in what fails when one unit is lost, so the choice should follow the redundancy requirement rather than the price per kilowatt.
- Deployment granularity: rack, row, or hall
- Blast radius when a single unit is unavailable
- Facility water availability at the intended location
- Ability to add capacity without disturbing live racks
The primary side is facility cooling
Everything on the primary side of the CDU belongs to the building: chilled water from a chiller, a condenser loop, a dry cooler, or a tower. The unit's cooling capacity is a function of what that facility cooling system can deliver and at what temperature, which is why a CDU quoted against cold chilled water can disappoint when connected to a warm loop. Confirm the primary supply and return temperatures, the available flow, and the pressure at the connection point before capacity is agreed, and check what the unit does when facility water drifts outside that window.
What the unit does, and does not do, for energy efficiency
A coolant distribution unit does not save energy by itself. What it does is make savings possible, by allowing the loop to run warm. Water at 40 degrees still cools a processor perfectly well, and a warm loop lets the site reject heat with dry coolers or a tower for most of the year instead of running a chiller. The savings live in that heat rejection choice, not in the unit. Two things move in the opposite direction and should be counted honestly: the pumps in the loop draw power that air cooling did not need, and every degree of temperature lost across the heat exchanger has to be made up somewhere upstream. Net energy use falls at high rack densities mainly because the server fans stop, which is a real saving of roughly a tenth of server power, and because heat removal by water is far more effective per watt spent moving it.
- Ask for net capacity with pump power included, not gross heat removal
- Establish the warmest supply temperature the servers are approved for
- Model the hours of the year that temperature lets you avoid the chiller
- Count the fan power the servers no longer draw as part of the result
Why GPU racks changed the requirement
Modern data centers deploying GPUs for AI training concentrate loads that traditional air cooling cannot reach, and the cooling loop has become part of the compute specification rather than a facility afterthought. That has pushed CDU capacity from hundreds of kilowatts toward multi-megawatt plants and made cooling performance a scheduling risk for the compute deployment. It also shifts where energy consumption sits: fan power falls inside the rack, pumping power appears in the loop, and the net result depends on how warm the loop is allowed to run.
Reliability in a high-performance computing hall
In an air-cooled data center a cooling failure gives operators minutes of thermal inertia. With direct-to-chip cooling, the heat load per rack is high and the fluid volume is small, so loss of coolant flow becomes critical very quickly. That changes what reliability means in the specification: pumps, controllers, valves, and power feeds all need a defined behavior on failure, and the sequence that protects the compute has to be demonstrated during commissioning rather than described in a datasheet.
Benefits of a coolant distribution unit in liquid cooling
The practical benefit of a coolant distribution unit is control at the boundary between two water systems. It keeps the clean technology loop away from facility water, gives data center operators one place to filter coolant and detect leaks, and holds the server supply temperature steady when either the compute load or the building loop changes. That makes phased deployment possible: a site can add liquid-cooled rows without redesigning the entire facility cooling system each time. It also creates a measurable heat-transfer boundary for commissioning and energy accounting. None of those benefits is automatic, because an undersized heat exchanger, weak controls, or poor water chemistry simply concentrates the failure in one machine. The specification has to prove the benefit at the site's actual duty point.
Service and spares
Price response time, on-site spares, pump and seal replacement intervals, and the availability of a like-for-like unit years from now. A CDU is a long-lived asset attached to a fast-moving server roadmap.