Liquid cooling is not one product. It is a chain of interfaces from the silicon package to the outdoor heat-rejection equipment. The correct architecture depends on heat flux, rack density, facility temperatures, service model, water constraints, and deployment speed.
Start with the heat path
Define which components move to liquid, what remains on air, and where each thermal boundary sits. A direct-to-chip system usually includes cold plates, hoses, quick-disconnect couplings, rack manifolds, the clean loop that serves the servers, a coolant distribution unit, and the building's own water loop. Immersion changes the hardware and service boundary by placing most or all IT components inside a dielectric bath.
- Processor and rack design heat loads
- Liquid capture ratio and residual air load
- Technology-loop supply and return temperatures
- Facility approach temperature and design-day heat rejection
Normalize supplier ratings
A megawatt label is not enough. Capacity changes with coolant temperatures, flow, pressure, fluid properties, redundancy state, and the allowed approach across heat exchangers. Compare vendors at one project duty point.
- Net capacity with one redundant component unavailable
- Pumping power and pressure drop at the selected duty
- Water quality and wetted-material requirements
- Controls, alarms, telemetry, and fail-safe behavior
Design the operating model with the equipment
Liquid systems change commissioning, leak response, maintenance, spare parts, fluid testing, and server service. Include operations teams before freezing the mechanical design.
Choose the architecture from the fleet
Start with the servers that will actually arrive, not with a preferred cooling product. Factory-integrated direct-to-chip racks preserve familiar rack and service patterns and have the broadest support for current accelerator systems. Rear-door heat exchangers suit mixed fleets because they leave servers untouched. Immersion can remove server fans and most room air, but it requires hardware and technicians to be organized around tanks. A mixed campus may use all three. The decision record should name which equipment is approved, which heat remains on air, and who owns each interface.
- Approved server and accelerator configurations for the first three deployment waves
- Heat captured to liquid at full rack load, not processor load alone
- Hardware warranty position for cold plates, fluids, and submerged parts
- Migration path when the server generation changes
Make temperature the common design language
Temperature connects the server, coolant distribution unit, and outdoor plant. The compute vendor defines the warmest liquid it will accept. The distribution equipment adds an approach temperature between loops. Weather determines how often the site can reject heat without refrigeration. Put those values on one diagram. A supplier capacity rating measured with colder water should be recalculated at the project's design point. Warmer operation can reduce compressor hours and improve heat reuse, but only when every component in the loop is approved for it.
Size the residual air system
Liquid cooling rarely captures every watt. Memory, storage, optics, power supplies, and parts of the network may still reject heat into the room. Ask the server vendor for a liquid capture ratio at the exact configuration, then convert the remainder into room and row air loads. This prevents two common errors: removing too much air-handling capacity and buying a liquid system whose headline rack rating ignores the heat left behind. Treat hybrid operation as a permanent state, including controls and failure response.
Compare economics at the system boundary
A useful cost model includes server premiums, cold plates or tanks, manifolds, distribution units, facility pipework, heat rejection, residual air equipment, controls, fluid, commissioning, and service. It also credits server fan power that disappears, compressor hours avoided, and capacity recovered from the hall. Compare the same load, climate, redundancy, and useful life. A component price or power usage effectiveness claim cannot settle the decision because it may exclude most of the thermal chain.
Commission the failure cases
Factory tests prove that equipment can carry a stated load. Site acceptance must prove that the assembled system protects compute when something fails. Test loss of one pump, controller, power feed, heat exchanger, communication path, and facility-water supply. Verify leak alarms, automatic isolation, server throttling, and restart. Record fluid cleanliness before connecting servers and again after the loop has circulated. Handover is complete only when operators can perform the response without the commissioning team.
Use a stage gate for procurement
The first gate proves technical fit with one rack and one duty point. The second proves the facility interface and operating procedure in a small pod. The third releases repeat orders only after measured performance and failure testing. This keeps a pilot from becoming a campus standard by momentum alone. It also gives suppliers clear evidence requirements: approved hardware, net capacity, controls integration, service coverage, and repeatable commissioning documents.