xAI · AI model development · United States (Memphis, Tennessee)
xAI Colossus: Spectrum-X Ethernet for a 100,000-GPU training cluster
xAI's Colossus cluster in Memphis started with 100,000 NVIDIA Hopper GPUs, connected with NVIDIA Spectrum-X Ethernet (SN5600 switches and BlueField-3 SuperNICs) to train Grok models. NVIDIA reports 95 percent data throughput, compared with the 60 percent it attributes to standard Ethernet at this scale.
- Status
- In production
- Status as of
- 28 Oct 2024
- NVIDIA technologies
- NVIDIA Spectrum-X Ethernet, NVIDIA BlueField
The challenge
Training one model across 100,000 GPUs needs a network where many large flows run at once without collisions. NVIDIA says standard Ethernet at this scale produces thousands of flow collisions and reaches only 60 percent of possible data throughput.
Data Center Frontier reads the 19-day GPU deployment as a sign of xAI's focus on scaling its infrastructure quickly; xAI's own goals are not stated in the cited sources.12
What was implemented
Colossus started with 100,000 NVIDIA Hopper GPUs in Memphis, Tennessee, in a former Electrolux factory according to Data Center Frontier, which also says Dell Technologies and Supermicro partnered with xAI on the build. Its RDMA network uses NVIDIA Spectrum-X Ethernet: Spectrum SN5600 switches (Spectrum-4 ASIC, ports up to 800 Gb/s) paired with BlueField-3 SuperNICs.
Data Center Frontier reports liquid-cooled Supermicro servers with eight GPUs each, eight servers per rack, and a 400 GbE network interface per GPU plus one more per server. At the time of NVIDIA's release xAI was doubling the cluster toward 200,000 Hopper GPUs.
The same article notes cost, power and cooling concerns for a system of this size; the Tennessee Valley Authority approved more than 100 MW of power for the site.12
Reported outcomes
Network keeps 95 percent data throughput with Spectrum-X congestion control1
95% vs 60% (NVIDIA's figure for standard Ethernet)
MeasuredNo application latency degradation or packet loss from flow collisions across three network tiers1
Zero collision-related packet loss
MeasuredCluster built in 122 days, with training starting 19 days after the first rack arrived12
122 days build; 19 days to first training
Measured
Each outcome links to its source; the label there says whether the company, NVIDIA or a third party reported it.
What others can learn, and the limits
What others can learn: at very large GPU counts the network, and its congestion control, decides how much of the GPU spend turns into training work, so it should be designed with the cluster rather than added later. Limits: very few organizations build at this scale, and few buyers can secure this much hardware at once. Every performance figure comes from NVIDIA; the 60 percent baseline is NVIDIA's own description of standard Ethernet, not a measured comparison on this site, and no test method is published. Clusters of a few hundred GPUs often run well on standard Ethernet or InfiniBand fabrics without these features.
Explore a similar project
Use the AI Factory Efficiency Lab with your own assumptions. Results are independent of this case.
Sources
Thank you. Your correction was sent.
The editors check it against the sources. If you left an email address, they may reply about it.