ALPHA HARDWARE
Empire AI: Alpha Hardware Specifications
A full inventory of the compute, networking, and storage hardware behind the Alpha cluster.
Compute — 28 GPU nodes + 1 CPU node, 224 GPUs
24 HGX nodes
- Nodes
alphagpu01–18— 8× H100 80GB GPUs per node - Nodes
alphagpu19–24— 8× H200 141GB GPUs per node
- 10× 400Gb/s ConnectX-7 NIC cards per node (8 for InfiniBand, 2 for Ethernet)
- 30TB NVMe caching space per node
- 2TB of system memory per node
4 RTX PRO 6000 Blackwell nodes
- Nodes
alphagpu51–54— 8× RTX PRO 6000 Blackwell ~96GB GPUs per node (32 total) - Part of the same general-access
alphapartition as the H100/H200 nodes, reserved for single-GPU workloads — see the Job Submission and QoS Overview article for routing behavior
1 CPU node
- Node
alphacpu01— 2× Intel Xeon Platinum 8568Y+ (96 cores total, x86_64), ~1TB RAM - Part of the dedicated
cpupartition for CPU-only workloads
Grace-Grace
Grace-Grace shares the same BCM/Slurm instance as Alpha and is reachable via the grace partition.
60 Grace CPU nodes
- Nodes
betagg01–60— 2× NVIDIA Grace (Neoverse-V2) sockets per node, 144 cores total, ARM (aarch64), ~480GB RAM per node - CPU-only, part of the
gracepartition
1 Grace Hopper node
- Node
alphagh01— NVIDIA GH200 Grace Hopper Superchip: 1× Grace CPU (72 cores, ~573GB system memory) + 1× Hopper GPU (~96GB HBM3e) - Part of the
gracepartition, alongside the 60 Grace CPU nodes above
Networking
Non-blocking NDR fabric, cabled for rail configuration
- 8 network switches, 96 optical connections
Service & access nodes
2 service nodes
- 1 login nodes
- 1 cluster management nodes (NVIDIA Base Command, licensed for all gear)
2 data transfer nodes
Storage — ~30PB total
~10PB of DDN storage (Lustre/EXAScaler)
- Working/scratch tier for active training checkpoints
~20PB of VAST storage
- Home directories (100GB/user hard quota) + project directories
Was this article helpful?
That’s Great!
Thank you for your feedback
Sorry! We couldn't be helpful
Thank you for your feedback
Feedback sent
We appreciate your effort and will try to fix the article