News
News

Rack-Level CDU Sizing for AI GPU Clusters: H100, B200, and GB200 NVL72

Sep. 10, 2026

A single NVIDIA H100 SXM pulls 700 W, a B200 SXM pulls 1,000 W, and a GB200 NVL72 rack holds 72 of them for a measured 120–130 kW per rack — so a 1 MW AI cluster is not "a bigger version of an enterprise data center." The plumbing has to be designed for an order-of-magnitude jump in rack density, and the unit that decides whether the plumbing works is the rack-level coolant distribution unit (CDU). Icicleflow supplies rack CDUs from 35 kW per cabinet to 300 kW facility-class units, sized for AI training fleets running direct-to-chip cold plate loops on H100, B200, and GB200 platforms. Sizing a CDU by nameplate IT load is the most common mistake; the right answer is sized for the actual secondary loop flow, ΔT, and redundancy class the GPU platform requires.

Summary. This article explains how to size a rack-level CDU for an AI GPU cluster, maps each NVIDIA platform to the correct Icicleflow CDU class, lists the procurement items to verify before issuing a purchase order, and walks through a representative GB200 NVL72 deployment handled by Icicleflow.


Why AI GPU Clusters Are a Different Plumbing Problem

  • An H100 8-way node holds 8 GPUs at 700 W each, or 5.6 kW of GPU TDP per node, plus 1–2 kW of CPU and memory. A standard 42U rack at 8–10 nodes pulls 50–70 kW, which is past the limit of a single cold plate loop without a rack CDU.

  • A B200 8-way node doubles the GPU TDP to roughly 8 kW, so the same rack now pulls 80–100 kW. The secondary loop flow rate, pump head, and heat exchanger size all scale with the GPU TDP, not with the rack nameplate.

  • A GB200 NVL72 rack integrates compute and switch trays into a single rack for 120–130 kW. The CDU for this class is closer to a small facility CDU than to a server-room piece of hardware, which is why the Icicleflow rack and facility CDU family covers the full range from 35 kW rack units to 300 kW cabinet units.


How to Size a Rack CDU for an AI Cluster

Step 1 — Decide which GPU platform the loop serves

  • H100 SXM 8-way (HGX H100): 5.6 kW GPU TDP per node, 50–70 kW per rack, served by a 35 kW class rack CDU with two CDUs for N+1 redundancy.

  • H200 SXM 8-way: identical plumbing to H100, with a modest increase in memory power; no CDU change needed.

  • B100 / B200 SXM 8-way (HGX B200): 8 kW GPU TDP per node, 80–100 kW per rack, served by a 70 kW class rack CDU per pair of racks, or by a 100 kW class CDU per rack.

  • GB200 NVL72: 120–130 kW per rack, served by a single facility-class 300 kW CDU with redundant pumps, dual power, and a plate-frame heat exchanger to the campus loop.


Step 2 — Size the secondary loop flow and ΔT

  • A 100 kW rack needs roughly 3–4 L/s of secondary loop flow at a 6–8 °C ΔT to stay within the GPU cold plate specification. Reducing the flow to 2 L/s forces the ΔT up to 12 °C, which pushes the cold plate outlet temperature past the GPU's thermal limit.

  • The facility water temperature sets the lower bound on the secondary loop inlet. At 18 °C facility inlet, a 6 °C ΔT gives a 24 °C secondary outlet, well within spec. At 25 °C facility inlet, the secondary outlet hits 31 °C, and the cold plate specification has to be re-checked.

  • The pump head has to be specified for the actual loop, including the cold plate, the quick disconnects, the rack manifold, and the CDU's internal heat exchanger. A pump that passes a bench test at zero flow will not push a 12-rack manifold.


Step 3 — Specify the redundancy class

  • N means a single CDU. A single failure takes the rack offline. Acceptable for non-production inference.

  • N+1 means one redundant CDU per cluster. A failure can be switched over without a workload interruption. Standard for production AI training.

  • 2N means two independent CDU chains, each capable of carrying the full load. Required for some hyperscaler financial-services deployments and for very large GB200 NVL72 rows.

  • Icicleflow's rack CDUs ship with dual pumps as standard, with automatic failover and a documented service contract.


Step 4 — Specify the telemetry interface

  • AI training jobs run for weeks, and an anomaly in coolant temperature, pressure, or flow rate is the earliest signal of a pump or cold plate problem. A CDU that exports through Redfish, IPMI, or a documented REST API lets the IT monitoring stack own alerting.

  • A CDU that only has a local display becomes invisible at 2 a.m., which is exactly when it fails. This is the single most-cited complaint from operators running 24/7 AI training fleets.


Where the Right CDU Class Changes by Platform

GPU Platform

Rack Power

Recommended CDU Class

Redundancy

Coolant

Common Watch-Out

HGX H100 8-way, 8-node rack

50–70 kW

Icicleflow 35 kW CDU per pair of racks

N+1

Deionized water + inhibitor

Pump head for 12-rack loop

HGX H200 8-way, 8-node rack

50–75 kW

Icicleflow 35 kW CDU per pair of racks

N+1

Deionized water + inhibitor

Same as H100

HGX B200 8-way, 6-node rack

70–90 kW

Icicleflow 70 kW class CDU per rack

N+1

Deionized water + inhibitor

Secondary ΔT at high load

HGX B200 8-way, 8-node rack

90–100 kW

Icicleflow 100 kW class CDU per rack

2N for production

Deionized water + inhibitor

Facility water at 25 °C inlet

GB200 NVL72 rack

120–130 kW

Icicleflow 300 kW cabinet CDU

2N for hyperscaler

Deionized water + inhibitor

Plate-frame heat exchanger sizing

GB200 NVL72 row (multiple racks)

500–1,000 kW per row

Multiple 300 kW CDUs in parallel

2N

Deionized water + inhibitor

Makeup water and leak detection

B200 liquid-cooled PCIe cards in edge site

15–25 kW

Icicleflow 35 kW CDU for the rack

N

Deionized water + inhibitor

High-ambient derating


What Buyers Report After Volume Deployment

What consistently works

  • Sizing the CDU for the secondary loop ΔT, not the rack nameplate. A 100 kW rack with a 6 °C ΔT and a 35 kW CDU will fail; a 100 kW rack with a 6 °C ΔT and a 100 kW CDU will not.

  • Specifying the pump head at the actual loop length. Operators who documented the loop length and the quick disconnect count before quoting the pump saw fewer cold plate ΔT surprises at full load.

  • Pairing the CDU telemetry with the GPU monitoring stack. A CDU that exports through Redfish cuts the mean time to detect from hours to minutes.

What consistently does not

  • Sizing the CDU at nameplate IT load with no headroom. A training job's peak is well above the nameplate TDP, and a CDU running at 95% capacity has no margin for pump ageing.

  • Treating GB200 NVL72 as a 100 kW deployment. The 120–130 kW measurement at full load leaves no room for an N+1 redundancy calculation inside a single 100 kW CDU.

  • Specifying facility water at 18 °C when the building plant delivers 25 °C in summer. The CDU has to be qualified at the worst-case inlet, not the design-day inlet.


Procurement Specification Checklist

Item

What to Verify

Why It Matters

Cooling capacity

Continuous at the worst-case inlet water temperature

Peaks are higher than nameplate

Secondary flow rate

Documented at the design ΔT, typically 6–8 °C

Determines the cold plate outlet temperature

Pump head

At the actual loop length, including quick disconnects

Long loops need higher head

Redundancy class

N, N+1, or 2N, with documented failover time

Determines workload continuity on failure

Heat exchanger

Plate-frame or brazed, sized for full load at worst-case inlet

Affects facility water coupling

Power input

Dual power supply standard for AI racks

Single PSU is a single point of failure

Telemetry interface

Redfish, IPMI, or documented REST API

Owns alerting through the IT stack

Leak detection

In-line conductivity plus spot sensors under the rack

Catches both fast and slow leaks

Coolant and service

Single fluid, factory-prefilled; spares stocked regionally

Prevents contamination and shortens MTTR


Application Case: GB200 NVL72 Deployment

A representative profile, drawn from a typical hyperscaler GB200 NVL72 deployment, illustrates how the four sizing steps interact. A hyperscaler was deploying a 2 MW GB200 NVL72 training cluster in a newly built hall, with a facility water plant designed for 18 °C inlet at full load.

The CDU selection was driven by three decisions. The 120–130 kW per rack measurement placed each rack above the 100 kW class, so the Icicleflow facility-class cabinet CDU was selected as the per-rack primary, with a second unit per rack for 2N redundancy. The secondary loop ΔT was specified at 6 °C to keep the cold plate outlet within the GB200 specification, which fixed the secondary flow at roughly 5 L/s per rack. The facility water coupling was sized with a plate-frame heat exchanger inside the cabinet CDU, sized for the full 130 kW at 18 °C inlet and verified at 25 °C inlet for the summer derate.

The qualification covered pump head at the actual rack loop length, failover time on simulated pump failure, and a 168-hour burn-in at full GB200 load. The CDU cleared all three on the first pass, and the customer attributed the result to having specified the secondary ΔT and the redundancy class in the same document as the cooling capacity. The deployment went live without a single CDU-related delay, and the operations team has since standardized on the same CDU class for B200 and H200 deployments in the same hall.


Professional Advice from Icicleflow

  1. Size the CDU for the secondary ΔT, not the rack nameplate. A 100 kW rack with a 6 °C ΔT and a 35 kW CDU will fail at full load; the same rack with a 100 kW CDU will not.

  2. Specify the redundancy class at the same time as the capacity. N+1 is the right default for production AI training, and 2N is required for some hyperscaler financial-services deployments.

  3. Qualify the CDU at the worst-case facility inlet temperature. A CDU that passes at 18 °C inlet may throttle at 25 °C inlet, which is exactly the temperature the building plant delivers in summer.

  4. Pair the CDU telemetry with the GPU monitoring stack. A CDU that lives outside the IT monitoring stack becomes invisible at 2 a.m., which is exactly when it fails.


Frequently Asked Questions

Q: How many kW of CDU capacity do I need for an H100 8-node rack?A: Plan for 50–70 kW per rack, with two 35 kW class CDUs for N+1 redundancy. The Icicleflow rack CDU family is the standard platform for this class.

Q: Can one facility-class cabinet CDU serve multiple GB200 NVL72 racks?A: For multiple racks, the cabinet CDU can serve as the facility-side aggregation, with smaller 35 kW or 70 kW CDUs as rack-level primaries. Topology depends on the redundancy class.

Q: What facility water temperature should I design for?A: For ASHRAE W-class cold plate loops, 18 °C inlet is the design point and 25 °C inlet is the summer derate. The CDU has to be qualified at both.

Q: Does the CDU need a separate leak detection system?A: Yes. In-line conductivity detection on the secondary loop plus spot sensors under the rack catches both fast leaks and slow seeps. A drip-only system misses the slow leak that wicks into cable insulation.

Q: Can Icicleflow build a custom CDU for a non-standard power or voltage?A: Yes. Icicleflow operates as a real ODM factory, with in-house brazing, structure design, and pump development. Custom CDUs typically run 8–12 weeks from drawing approval to first unit.


Talk to Icicleflow About Your AI GPU Cluster CDU Specification

If you are deploying an H100 training fleet, a B200 inference cluster, or a GB200 NVL72 row, Icicleflow can supply the rack CDU, the brazed heat exchanger, and the cold plate loop from one ODM engineering team. Send your GPU platform, target rack power, facility inlet water temperature, and deployment date to sales@icicleflow.com or use our contact page. We respond with a written specification and a sample plan within two business days.



Related News
All News
WeChat
0769-82078456