MENU

Blueprint for 100 kW+ GPU Cluster AI Data Centers Released: Liquid Cooling Becomes Indispensable

Soeteck USA
Overview
A practical blueprint for designing AI data centers to support high-density GPU clusters, exceeding 100 kW per rack with NVIDIA GB300-class deployments, has been unveiled. The report underscores that traditional data center assumptions are obsolete, making liquid cooling—particularly cold-plate technology, which accounts for about 80% of the liquid-cooled market—essential. This design shift necessitates significant adjustments in power conditioning, heat rejection, and grey space allocation for cooling and electrical systems. This will accelerate the construction of next-generation data centers capable of meeting AI’s computational demands.
In Depth

Key Findings

A specific blueprint for designing AI data centers capable of supporting ultra-high-density GPU clusters, such as NVIDIA’s GB300-class deployments exceeding 100 kW per rack, has been released. This design unequivocally demonstrates that traditional data center assumptions are no longer viable in the age of AI, positioning liquid cooling technology as an indispensable foundation. Cold-plate technology, currently comprising approximately 80% of the liquid cooling market, plays a central role. Consequently, significant re-evaluations are required in power conditioning, heat rejection, and the allocation of ‘grey space’ for cooling and electrical systems. This development provides crucial guidance for building next-generation data centers that meet escalating AI computational demands.

Technical / Clinical Details

AI data center design addresses several critical technical challenges and proposes corresponding solutions:

  • Ultra-High-Density Power Delivery: Traditional power distribution architectures are insufficient for supporting rack densities exceeding 100kW. For example, a transition from 48V to 54V bus voltages and the integration of high-efficiency Power Supply Units (PSUs) at the rack level are essential. Designs must prioritize minimizing power loss and ensuring stable power delivery.
  • Mandatory Liquid Cooling Technology: To efficiently remove the substantial heat generated by high-performance GPUs (e.g., GB200 GPUs can generate several kilowatts of heat), cold-plate liquid cooling is widely adopted. Direct contact between the coolant and the chip dramatically improves thermal transfer efficiency, keeping chip temperatures low. Critical design considerations include coolant management, leak prevention, and robust supply/return systems.
  • Rethinking Heat Rejection and Grey Space: Heat absorbed by liquid cooling is processed by external systems like chiller units or cooling towers. This necessitates optimizing the design and placement of ‘grey space’ (non-IT areas housing cooling and power equipment) within the data center facility for efficient heat dissipation and power provision. Grey space tends to occupy a larger proportion compared to traditional data centers.
  • Modular Infrastructure and Scalability: To accommodate rapidly evolving AI hardware, the modularization of power, cooling, and network infrastructure is being promoted. This enables the construction of flexible data centers that can quickly scale up or down according to demand.

These technical elements collectively define a new typology of data centers specifically optimized for AI’s computational needs.

Background & Context

The advent of generative AI and large language models (LLMs) has led to a dramatic leap in AI chip performance, consequently pushing power consumption and heat generation to unprecedented levels. High-performance GPUs like NVIDIA’s H100 and GB200 impose entirely different infrastructure requirements compared to traditional CPU-based servers, which existing data center designs cannot meet. As a result, the entire data center industry is urgently seeking new design principles and technologies optimized for AI workloads. Amidst concerns about data center shortages and power deficits, the construction of efficient and sustainable AI data centers is critically important for the growth of the entire AI ecosystem.

Strategic Significance & Outlook

With the establishment of this blueprint for AI data center design, the construction of data centers specialized in high-density AI computing is expected to accelerate over the next few years. Liquid cooling technology will become the de facto standard for data center cooling, with its performance and reliability continuing to improve. Furthermore, diversification of power sources, particularly the integration of renewable energy and new solutions like Small Modular Reactors (SMRs), may advance. Optimizing AI data centers will maximize the potential of AI technology, serving as a foundation for accelerating innovation in various fields such such as autonomous driving, robotics, scientific computing, and medical AI. This shift in design will also have long-term implications for data center location selection, energy strategy, and supply chains.

Source: https://soeteck.com/en/news-and-insights/blogs/ai-data-center-design/

Get our weekly technology intelligence — free

Receive an infographic that lets you judge at a glance whether each field’s analysis report is worth reading.

Subscribe Free — Weekly Tech Intelligence

By subscribing, you’ll receive Troy-Technical’s weekly technology intelligence newsletter.

  • Your email and selected fields are used only to deliver the newsletter.
  • We never share your information with third parties.
  • You can unsubscribe anytime via the link in each email.

See our Privacy Policy for details.

Takes about a minute · Unsubscribe anytime

Let's share this post !

Author of this article

Comments

To comment

TOC