Throughput · NVFP4
$80/hr $0.0222/sec · $58,400/mo at 24/7
Best tokens/$ for large MoE models. Large MoE throughput kings.
VRAM pool
180GB ×8 = 1.4TB
Position
vs A100 8×: ~3× throughput
Compute die
NVIDIA B200 180GB — NVFP4 tensor cores, tuned for large moe throughput kings.
HBM stack
180GB ×8 = 1.4TB of high-bandwidth attached memory — the whole model fits, no offload.
Interposer / NVLink
NVLink-C2C fabric so the 8 GPUs act like one pool.
Per-second metering
$80/hr → 0.0222/sec. Pause the deployment and the meter hits $0.
Scale-to-zero
Autoscale + region pins (GLOBAL / US / EU / APAC) on the control plane.