Rackscale AI Accelerators
- Exponential Industry
- Technology Area
- 14 Milestones
Fully co-designed rack- and POD-scale AI systems that combine accelerators, host CPUs, scale-up fabric (NVLink-class domains), and scale-out networking in one thermal/mechanical package. Lineage runs from multi-GPU HGX/DGX boxes through Grace Hopper NVL32 domains to liquid-cooled 72-GPU racks (GB200, GB300, Vera Rubin) and competing racks such as AMD Helios. NVIDIA, AMD, and OEM rack builders (Dell, Supermicro) ship the current generation for AI-factory halls.
Exponential Industry · Hidden Factory
Rackscale AI Accelerators
OpenAI Trains GPT-6 Astra on 100K+ NVIDIA Grace Blackwell NVLink72 at Abilene Campus
OpenAI operationalized frontier model GPT-6 Astra, trained across more than 100,000 liquid-cooled NVIDIA Grace Blackwell GB200 NVLink72 GPUs at Crusoe's 1.2 GW Abilene, Texas clean energy campus, marking a milestone in frontier AI scale /X (formerly Twitter)/.
SK hynix Breaks Ground on $4B Advanced Packaging Fab at Purdue Research Park
SK hynix held a ceremonial groundbreaking for its $4 billion advanced packaging fabrication and R&D hub at Purdue Research Park in West Lafayette, Indiana, establishing its first U.S. high-bandwidth memory (HBM) production facility /Purdue University News/.
NVIDIA Groq 3 LPX Enters Full Production; Nebius Deploys in Token Factory
NVIDIA announced full production of the rack-scale Groq 3 LPX system (256 LPUs, 128GB SRAM, 640TB/s bandwidth), with Nebius deploying clusters into the Nebius Token Factory for low-latency agentic decoding achieving 3,400 tok/s on Gemma 4 31B /NVIDIA Newsroom/ /Nebius/.
NVIDIA Expands NVLink Fusion with NVHBM and Announces Annapurna Labs Trainium4 Partnership
NVIDIA unveiled NVHBM custom high-bandwidth memory architecture for NVLink Fusion, moving the memory controller into the HBM base die, and announced Amazon's Annapurna Labs as the inaugural partner for AWS Trainium4 custom AI silicon /NVIDIA Blog/.
SpaceXAI adopts NVIDIA Vera CPU and Vera Rubin for Agentic AI
SpaceXAI announced deployment of NVIDIA Vera CPUs featuring 88 Olympus cores and 1.2TB/s memory bandwidth to accelerate orchestration and code execution for Grok agentic AI workloads /NVIDIA Newsroom/ /X (formerly Twitter)/.
First production NVIDIA Vera Rubin systems arrive at Microsoft datacenters
Satya Nadella (Aug 21, 2026) posted that it was delivery day at Microsoft datacenters as the first production Vera Rubins arrived, thanking NVIDIA and Azure hardware and datacenter teams /X/.
SK hynix publishes CPO roadmap in Nature Electronics
SK hynix newsroom Aug 20, 2026: paper "Co-packaged optics for high-performance computing and artificial intelligence" in Nature Electronics (s41928-026-01681-6), corresponding authors Seunghoon Hong (SK hynix) and Kyusang Lee (UVA). Targets >100 Tb/s per node, <1 pJ/bit, <10 ns latency; optics-centric photonic interposer between XPU and memory pools to break the rack/pod bandwidth wall /SK hynix/ /Nature Electronics/.
Nebius Group Issues $4.5B in Convertible Senior Notes for AI Data Center Expansion
Nebius Group issued $4.5 billion of convertible senior notes ($2.75B due 2030 and $1.75B due 2034) to fund greenfield AI data center construction and high-density GPU procurement across Europe and the United States /The Next Web/.
Etched delivers first inference rack to Jane Street
Etched PR Aug 18, 2026: first rack shipped last month; Jane Street said it tested the chip, is pleased with early results, and has its own rack running in its data center. Treated as the announcement day; ship month is July 2026 without a calendar day /GlobeNewswire/.
AMD Helios rack-scale AI launch
AMD launches Helios rack-scale AI infrastructure (72× MI455X) at Advancing AI 2026 (AMD blog).
First Dell PowerEdge XE9812 Vera Rubin NVL72 delivery to CoreWeave
Michael Dell announced delivery of the world's first working liquid-cooled Dell PowerEdge XE9812 NVIDIA Vera Rubin NVL72 server rack for CoreWeave (X, May 31, 2026).
NVIDIA Vera Rubin platform announcement
NVIDIA announces Vera Rubin platform and related racks at GTC; seven chips in full production.
NVIDIA Blackwell platform and GB200 NVL72 announcement (GTC 2024)
GTC 2024 (Mar 18) launch of the NVIDIA Blackwell platform: Blackwell GPU architecture, GB200 Grace Blackwell Superchip, liquid-cooled GB200 NVL72 rack-scale system (72 Blackwell GPUs + 36 Grace CPUs on fifth-gen NVLink), HGX B200, sixth-generation air-cooled DGX B200 (8 Blackwell GPUs, two 5th Gen Intel Xeon; up to 144 PFLOPS FP4), and broad cloud/OEM adoption commitments. Foundational rackscale-AI milestone that established the NVL72 rack product line later continued by GB300 and Vera Rubin NVL72; DGX B200 is the traditional air-cooled 8-GPU SuperPOD/BasePOD building block, not NVL72 /NVIDIA Newsroom/.
NVIDIA DGX GH200 AI supercomputer announcement
DGX GH200 large-memory AI supercomputer class with Grace Hopper Superchips and NVLink Switch System—predecessor NVLink domain architecture before Blackwell NVL72.
Frequently Asked Questions
What is Rackscale AI Accelerators and what industrial engineering problems does it address?
Fully co-designed rack- and POD-scale AI systems that combine accelerators, host CPUs, scale-up fabric (NVLink-class domains), and scale-out networking in one thermal/mechanical package. Lineage runs from multi-GPU HGX/DGX boxes through Grace Hopper NVL32 domains to liquid-cooled 72-GPU racks (GB200, GB300, Vera Rubin) and competing racks such as AMD Helios. NVIDIA, AMD, and OEM rack builders (Dell, Supermicro) ship the current generation for AI-factory halls.
What core technologies and manufacturing architectures comprise Rackscale AI Accelerators?
Key technical architectures and manufacturing innovations include:
- NVIDIA Vera Rubin Platform: Full-stack agentic AI infrastructure platform: Vera CPU, Rubin GPU, NVLink 6, ConnectX-9, BlueField-4, Spectrum-6, and Groq 3 LPU integrated into rack types (NVL72, Vera CPU, LPX, STX storage, SPX Ethernet).
- NVIDIA NVL72 Rack Product Family: Product family for NVIDIA liquid-cooled 72-GPU NVLink rack systems: GB200 NVL72 (Blackwell, Mar 2024), GB300 NVL72, Vera Rubin NVL72, and OEM form factors (e.g. Dell PowerEdge XE9812).
- NVIDIA Vera Rubin NVL72: Rack integrating 72 Rubin GPUs and 36 Vera CPUs via NVLink 6 with ConnectX-9 and BlueField-4. NVIDIA claims up to 10x higher inference throughput per watt vs Blackwell generation at one-tenth cost per token (Vera Rubin PR).
- NVIDIA NVLink Fusion: Rack-scale platform enabling customers to develop semi-custom AI infrastructure using the NVIDIA NVLink ecosystem. Partners (e.g. Marvell) supply custom XPUs and NVLink Fusion-compatible scale-up networking; NVIDIA supplies Vera CPU, ConnectX NICs, BlueField DPUs, NVLink interconnect, Spectrum-X switches, and rack-scale AI compute.
- Etched frontier inference cluster: Rack-scale inference cluster using Low Voltage Inference (LVI) for compute density at the same power and Cluster Scale Memory (CSM), a shared memory pool across the cluster. First customer Jane Street; >$1B in contracts claimed.
- AMD Helios: Open rack-scale AI platform co-designed around 72 AMD Instinct MI455X GPUs, 6th Gen EPYC "Venice" CPUs, Pensando networking / UALoE fabric, and ROCm software. Claims 2.9 EF dense FP4, 31 TB HBM4, 260 TB/s scale-up bandwidth (AMD Helios blog, Jul 2026).
Which companies and industrial facilities lead deployment in Rackscale AI Accelerators?
Leading industrial manufacturers, hyperscalers, and engineering operators include:
- NVIDIA Corporation: NVIDIA Expands NVLink Fusion with NVHBM and Announces Annapurna Labs Trainium4 Partnership (Aug 26, 2026) — NVIDIA NVLink Fusion expands with NVHBM custom high-bandwidth memory. By moving the memory controller into the HBM base die, NVHBM delivers up to 30% higher memory bandwidth compared to standard HBM4E, reduces HBM power by up to 15%, and frees up to 25% more silicon area on the XPU compute die.
- Microsoft Corporation: First production NVIDIA Vera Rubin systems arrive at Microsoft datacenters (Aug 21, 2026) — Delivery day at our Microsoft DCs as the first production Vera Rubins arrive. A huge thank you to our partners at @nvidia and our Azure hardware and datacenter teams for all the incredible work that brought us to this milestone!
- Advanced Micro Devices: AMD Helios rack-scale AI launch (Jul 23, 2026) — AMD launches Helios rack-scale AI infrastructure (72× MI455X) at Advancing AI 2026 (AMD blog).
- SK hynix: SK hynix Breaks Ground on $4B Advanced Packaging Fab at Purdue Research Park (Aug 27, 2026) — SK hynix Inc. held a ceremonial groundbreaking Thursday (Aug. 27) for its over $4 billion advanced packaging fabrication and R&D facility for AI memory in the Purdue Research Park, marking SK hynix's first U.S. high bandwidth memory production hub.
- Dell Technologies: First Dell PowerEdge XE9812 Vera Rubin NVL72 delivery to CoreWeave (May 31, 2026) — Michael Dell announced delivery of the world's first working liquid-cooled Dell PowerEdge XE9812 NVIDIA Vera Rubin NVL72 server rack for CoreWeave (X, May 31, 2026).
- Etched: Etched delivers first inference rack to Jane Street (Aug 18, 2026) — Etched shipped its first rack last month to Jane Street, and the quantitative trading firm is actively deploying the technology into its workloads.
What are key commercial projects and milestone achievements in Rackscale AI Accelerators?
Major industrial breakthroughs and commercial milestones include:
- NVIDIA Blackwell platform and GB200 NVL72 announcement (GTC 2024): NVIDIA today announced that the NVIDIA Blackwell platform has arrived — enabling organizations everywhere to build and run real-time generative AI on trillion-parameter large language models at up to 25x less cost and energy consumption than its predecessor.
- AMD Helios rack-scale AI launch (Jul 23, 2026): AMD launches Helios rack-scale AI infrastructure (72× MI455X) at Advancing AI 2026 (AMD blog).
- Etched delivers first inference rack to Jane Street (Aug 18, 2026): Etched shipped its first rack last month to Jane Street, and the quantitative trading firm is actively deploying the technology into its workloads.
- First production NVIDIA Vera Rubin systems arrive at Microsoft datacenters (Aug 21, 2026): Delivery day at our Microsoft DCs as the first production Vera Rubins arrive. A huge thank you to our partners at @nvidia and our Azure hardware and datacenter teams for all the incredible work that brought us to this milestone!
- NVIDIA Expands NVLink Fusion with NVHBM and Announces Annapurna Labs Trainium4 Partnership (Aug 26, 2026): NVIDIA NVLink Fusion expands with NVHBM custom high-bandwidth memory. By moving the memory controller into the HBM base die, NVHBM delivers up to 30% higher memory bandwidth compared to standard HBM4E, reduces HBM power by up to 15%, and frees up to 25% more silicon area on the XPU compute die.
- SK hynix Breaks Ground on $4B Advanced Packaging Fab at Purdue Research Park (Aug 27, 2026): SK hynix Inc. held a ceremonial groundbreaking Thursday (Aug. 27) for its over $4 billion advanced packaging fabrication and R&D facility for AI memory in the Purdue Research Park, marking SK hynix's first U.S. high bandwidth memory production hub.
Related Technology Ontologies
Physical AI & Embodied Robotics
Artificial intelligence systems and autonomous machines that operate directly in and interact with the physical world, bridging bits to atoms. Spanning foundation models (Vision-Language-Action and World Foundation Models), physics-based simulation with synthetic data (Omniverse, Isaac Sim) to safely bridge Sim2Real, and embedded runtime computers (Jetson Thor, DRIVE AGX) executing closed-loop perception-action loops across humanoids, AMRs, adaptive manipulators, autonomous mobility, and smart spaces.
Data Center Power & Behind-the-Meter Generation
Generation and storage sited at or next to AI/cloud campuses so the load does not wait on a transmission queue. Common stacks pair gas turbines or engines with on-site BESS; some designs add behind-the-meter solar. Hyperscalers and developers announced multi-hundred-megawatt gas+BESS campuses in 2025–2026 to serve rack-scale AI halls.
AI Inference Silicon
Dedicated application-specific processors and accelerators engineered for high-throughput, low-latency AI model inference in data centers and edge deployments. Architectures overcome the memory wall by utilizing massive on-chip SRAM, high-bandwidth memory (HBM3e), and specialized matrix processing engines. Commercial implementations include Etched Sohu ASICs, Groq Language Processing Units (LPUs), Tenstorrent Wormhole, and Cerebras CS-3 systems.