Проектирование центров обработки данных для искусственного интеллекта has shifted from a facility-engineering exercise into a compute-performance discipline. The global AI data center market is projected to grow from USD 236.4 billion in 2025 to USD 933.7 billion by 2030, a 31.6% compound annual growth rate, according to analysis published in August 2025. Rack densities that reached 2-10 kW in traditional facilities now exceed 100 kW in NVIDIA GB300-class deployments, with 2025 reference designs pointing toward 600 kW racks.
Why Does AI Data Center Design Differ From Conventional Facility Planning?
Conventional facilities were built around predictable, low-density IT. Servers drew 2-10 kW per rack, air cooling sufficed, and power chains were sized with generous margins. AI workloads break those assumptions. GPU clusters consume power in violent, workload-dependent bursts rather than steady loads, and they generate heat densities that air handlers cannot remove.
According to Omdia senior research director Vladimir Galabov, hyperscale deployments are experiencing a sharp increase in density that pushes facility infrastructure far beyond historical norms. A single 600 kW AI rack can consume more electricity in a couple of hours than an average family home uses in a month. Consequently, AI data center design now sizes power, cooling, networking, and layout from the GPU workload backward, not from generic building standards forward.
What Power Density Should an AI Data Center Design Target?
The target density determines every downstream decision, so it must be set from the actual GPU roadmap, not vendor marketing. AI data center design practice clusters around three density tiers. Traditional enterprise racks remain in the 10-30 kW range, and the Uptime Institute Global Data Center Survey 2025 notes that most facilities still stay below 30 kW per rack. Purpose-built AI factories standardize on 50-100 kW racks with direct-to-chip liquid cooling.
NVIDIA’s 2025 reference architectures push further, describing rack enclosures that support 500+ GPUs and draw 600 kW each, roughly five times the power of the highest-density racks in use today.
Choosing a density tier early matters because structural loading, busway sizing, cooling capacity, and floor layout all cascade from that number. Designers who plan for 30 kW and discover a 100 kW requirement mid-project face a retrofit far more costly than building right the first time. In a mature AI data center design, every rack position has a declared power ceiling, enforced through smart PDUs and capacity management software.

What Cooling Architecture Fits an AI Data Center Design?
Cooling is the first subsystem where traditional assumptions fail, and it is where AI data center design diverges most sharply from legacy practice. Air cooling can support modest increases, but the Uptime Institute survey shows average PUE has barely improved for six consecutive years, constrained partly by legacy infrastructure. High-density AI racks require liquid cooling. Cold-plate liquid cooling covers about 80% of the liquid-cooled market because it works with existing server architectures and offers the lowest total cost of ownership.
Immersion cooling pushes PUE down to approximately 1.05, and the immersion segment is projected to reach USD 5.8 billion by 2030 with a 39% CAGR. Regulatory pressure accelerates the shift: China requires new mega data centers built after 2025 to achieve PUE no higher than 1.25, which makes liquid cooling effectively mandatory. Direct-to-chip cooling is now standard in high-density AI deployments, often coexisting with air cooling for non-GPU equipment. Any AI data center design that relies on air alone for GPU halls will hit a thermal ceiling measured in megawatts.
What Power Delivery Architecture Supports AI Workloads Reliably?
Power delivery is the least visible but most expensive subsystem in AI data center design. GPU power draw is not a flat load; it oscillates rapidly as training steps begin and end. The power chain must absorb those swings without tripping protection or degrading grid stability. Traditional UPS systems achieve 90-94% efficiency, while 800V high-voltage DC architectures reach 97% or better and cut capital expenditure by roughly 20%. NVIDIA’s GB300 platform shows the rest of the story: energy-storage-enhanced power shelves reduced peak grid demand by 30% while training the Megatron LLM.
For your own facility, this means specifying power shelves with integrated battery buffering rather than oversized transformers. Backup battery units and supercapacitor packs are becoming standard at the rack level, with response times under one millisecond for the pulse-power compensation that GPU training demands. The electrical room must also be sized for the redundancy tier actually operated, not the Tier III badge claimed on paper. A resilient AI data center design treats power as a control problem, not a sizing problem.
What Network Topology Do GPU Clusters Demand?
AI data center design fails at the network layer more often than at the power layer, because distributed training has brutal bandwidth and latency requirements. Traditional leaf-and-spine networks tuned for north-south web traffic cannot sustain the east-west traffic patterns of collective communication. Training jobs move gradients, activations, and optimizer states between GPUs, and network saturation directly lengthens time-to-train. Modern GPU clusters separate scale-up and scale-out networks: scale-up connects GPUs within a node through high-bandwidth fabrics such as NVLink, while scale-out connects nodes through InfiniBand or 400G/800G Ethernet fabrics.
McKinsey estimates that data centers will require USD 6.7 trillion in cumulative capital expenditure by 2030 to meet compute demand, and a meaningful share of that budget goes to switching infrastructure. The rule: model communication patterns before buying switches, and never let network oversubscription become the hidden bottleneck. AI data center design must therefore treat network architecture as a first-class citizen, with oversubscription ratios set explicitly at every tier.
Why Does AI Data Center Design Prioritize Grey Space Over White Space?
Legacy data center layouts allocated most floor area to white space, the IT equipment halls, and squeezed cooling and electrical gear into whatever remained. AI data center design inverts that ratio. Dense GPU racks demand more power conditioning, more heat rejection, and more liquid distribution per square meter, so grey space for cooling and electrical systems now rivals or exceeds white space in purpose-built AI factories.
Practical implications: taller structural floors to support liquid loops, stronger slab loading for 600 kW racks, and modular power skids deployed incrementally. If grey space is undersized at the design stage, the facility hits a hard ceiling regardless of how many GPU racks the white space can physically hold. A common mistake in AI data center design is copying legacy floor plans and only swapping the cooling units, which leaves no room for the piping, pumping, and heat rejection that liquid cooling requires.
What Steps Should Be Followed When Designing an AI Data Center?
A disciplined sequence prevents the most expensive mistakes. First, define the target workload: training, inference, or both, because training is network- and storage-hungry while inference is latency-sensitive. Second, freeze the GPU roadmap and compute the rack density tier, power envelope, and cooling method for each hall. Third, size the electrical chain from the peak-with-buffering load and specify energy-storage power shelves. Fourth, design the network fabric around measured communication patterns, splitting scale-up from scale-out. Each of these decisions belongs in the early AI data center design phase, before concrete is poured.
Fifth, allocate grey space for liquid distribution, heat rejection, and electrical rooms before finalizing white space layouts. Sixth, validate against ASHRAE 90.4, which sets minimum efficiency requirements for cooling and power systems and now addresses water-cooled IT in its 2022 addenda. Seventh, plan for incremental deployment so early halls operate while later ones are built, a pattern increasingly common in AI data center design. Following this sequence keeps the whole project aligned with the compute roadmap and avoids the retrofit trap.
Which Metrics Validate an AI Data Center Design?
Design quality shows up in measurable indicators, and every AI data center design should be benchmarked against them before construction begins. PUE remains the headline efficiency metric, but the Uptime Institute survey found that one in ten outages still causes serious or severe disruption, so availability engineering deserves equal attention. PUE targets below 1.25 are feasible with liquid cooling but require the mechanical load component and electrical loss component of ASHRAE 90.4 to stay within climate-zone limits.
Rack utilization, measured as achieved compute throughput per dollar of facility capex, tells whether the design actually serves AI workloads. Time-to-power, the delay between rack installation and full GPU operation, exposes bottlenecks in busway, cooling, and network delivery.
Conclusion: How Should AI Data Center Design Evolve Next?
AI data center design is no longer a marginal optimization of an existing template; it is a structural break with three decades of facility practice. Density targets, cooling architecture, power delivery, network topology, and space allocation must all derive from the GPU workload, and the margin for error is measured in millions of dollars of stranded compute. The facilities that succeed lock density tiers early, standardize liquid cooling, buffer the power chain, and validate against ASHRAE and Uptime data. Start the AI data center design from the workload, not from the building, and the rest falls into place.

















