Calculate the complete power path
Use maximum sustained node power rather than an attractive average. Verify building feeds, UPS, switchboards, breakers, cable gauge, connectors and PDUs. In an A/B design, either branch must carry the load after the other fails. Include neighboring equipment, growth, protection selectivity, grounding and emergency shutdown.
Convert kilowatts into heat
Nearly all electrical input becomes heat. A conventional hot-aisle/cold-aisle room may not support a dense GPU rack even when total cooling capacity looks sufficient. Measure temperature at server inlets, airflow, pressure, blanking panels and recirculation. High density may require in-row or liquid cooling, adding CDU capacity, coolant quality, pump resilience, connections, leak detection and clear operational ownership.
Check mechanics and serviceability
Confirm usable rack depth, weight limits, floor loading, rail position and the delivery path. Leave enough service clearance to replace power, network and cooling components without disturbing adjacent systems. Align switch and server airflow, and keep cabling away from exhaust and sliding equipment.
Design network and storage with compute
An accelerator is idle when data cannot arrive. Separate production, management, storage and scale-out traffic; plan ports, optics, cable lengths and switch oversubscription. For training and RAG, evaluate sequential throughput, metadata behavior, concurrency and backup windows as well as raw storage capacity.
Accept the platform under combined load
Stress CPU, GPU, memory, network and disks at the same time. Record power per feed, inlet and exhaust temperatures, accelerator clocks, interface errors and cooling behavior after a component failure. The deliverable should be a rack map, A/B budget, thermal model, connection plan, monitoring thresholds and a list of facility work—not a vague statement that capacity is sufficient.

