Direct-to-chip cooling has become the default liquid architecture for many rack-scale AI systems because it targets GPUs and CPUs while preserving standard rack service patterns. The hard work is at the interfaces.
The four design boundaries
Treat the server loop, the rack distribution, the coolant distribution unit, and the building loop as four separate scopes of work. Each boundary needs an owner, an acceptance test, and a compatible materials list.
- Server cold plates, hoses, and quick disconnects
- Rack manifolds and branch isolation
- CDU pumps, heat exchangers, filters, and controls
- Facility water, chillers or economizers, and heat rejection
Capacity follows temperature
What a coolant distribution unit actually delivers depends on the water temperatures it is given and how closely the two loops approach each other. Warmer technology coolant can improve economizer hours and heat-reuse potential, but the server vendor defines the allowable inlet range.
Plan for mixed air and liquid loads
Cold plates rarely capture every watt. Quantify the air-side remainder for memory, networking, storage, power supplies, and rack leakage before downsizing room cooling.
Start with a qualified server configuration
The cold plate is part of the server. Its geometry, mounting pressure, materials, hoses, and firmware assumptions follow a specific processor and chassis. The lowest-risk route is a factory-integrated liquid-cooled system with a written warranty. A separately integrated plate can work, but the buyer then needs a named party to own contact quality, leak risk, and thermal performance. Record the exact server bill of materials and do not assume that approval transfers to the next accelerator generation.
Define the secondary loop
The technology cooling system carries fluid from the coolant distribution unit to rack manifolds and cold plates. Specify fluid chemistry, wetted materials, filtration, pressure, flow, hose and coupling standards, expansion volume, and cleanliness. Narrow cold-plate channels make construction debris and corrosion particles consequential. Flush and sample the loop before servers connect. The contractor, distribution-unit vendor, and server vendor should sign one compatibility schedule rather than provide three unrelated lists.
- Permitted fluid and inhibitor concentration
- Metals, elastomers, hoses, and quick disconnects in the wetted path
- Particle and chemistry limits at handover and during operation
- Maximum pressure and flow under every control state
Choose distribution by failure domain
In-rack units isolate faults to one cabinet but multiply pumps, filters, and assets. Row units serve a pod and balance deployment speed with maintenance count. Multi-megawatt units reduce asset count but place more load behind shared pumps, controls, and headers. Select the format from the acceptable blast radius and phasing plan, not from price per kilowatt. Ask each bidder to show the load carried with one component unavailable and the steps needed to isolate a live branch.
Connect to the facility at a real duty point
The primary loop may come from chillers, dry coolers, cooling towers, or a mixed plant. State its supply temperature, return target, pressure, water quality, and worst-day drift. The coolant distribution unit can only transfer heat when there is enough temperature difference across its heat exchanger. A two-megawatt nameplate tested with cold water may deliver less on a warm site loop. Require the capacity curve, pump power, and approach temperature at both peak and part load.
Control condensation and leaks differently
Condensation occurs when a surface falls below the room dew point. The usual defense is to run the technology loop warm and enforce a lower temperature limit in controls. Leak management uses sensors, drip paths, pressure behavior, automatic isolation, and non-conductive trays where needed. Do not combine the two into a vague water-risk statement. Commission dew-point response, one leaking branch, a failed sensor, and loss of communication as separate cases.
Write the service procedure before purchase
A technician needs to know how to isolate a rack, verify pressure is removed, disconnect a server, contain residual fluid, install a replacement, purge air, and return the branch to service. That procedure determines valve locations, hose length, access, and spare holdings. Ask the supplier to demonstrate it with the rack loaded as it will be in production. Time the task and state which warranty covers any fluid contact or coupling failure.
Acceptance evidence
A useful handover pack includes pressure and leak tests, flush and chemistry results, control-point verification, load tests at the specified temperatures, failure-state tests, server throttling behavior, and operator training records. Repeat key checks after the first months of operation because filters collect construction residue and fluid chemistry settles. Capacity alone is not acceptance; the system must be maintainable and observable under production constraints.