runaiicloud
ModelsGPUsPricingDocsConnectCompareEnterprisePlayground
Log inGet started
← All GPUs
DIEHBMINTERPOSERNVSWITCHSUBSTRATESUBSTRATE · POWER & THERMALINTERPOSER · NVLink-C2C FABRICHBMHBMHBMHBMB200NVFP4 · 8×NVSWITCH

Throughput · NVFP4

8× NVIDIA B200 180GB

$80/hr $0.0222/sec · $58,400/mo at 24/7

Best tokens/$ for large MoE models. Large MoE throughput kings.

VRAM pool

180GB ×8 = 1.4TB

Position

vs A100 8×: ~3× throughput

Deploy this →Configure sizeCompare providers →

What you're renting

Compute die

NVIDIA B200 180GB — NVFP4 tensor cores, tuned for large moe throughput kings.

HBM stack

180GB ×8 = 1.4TB of high-bandwidth attached memory — the whole model fits, no offload.

Interposer / NVLink

NVLink-C2C fabric so the 8 GPUs act like one pool.

Per-second metering

$80/hr → 0.0222/sec. Pause the deployment and the meter hits $0.

Scale-to-zero

Autoscale + region pins (GLOBAL / US / EU / APAC) on the control plane.

Reserve it in minutesDeploy →

Other presets

8× NVIDIA B300 288GB

$96/hr · FP8

Max quality for frontier dense models

4× NVIDIA B300 288GB

$48/hr · NVFP4

Smallest frontier-capable footprint

4× NVIDIA B200 180GB

$40/hr · NVFP4

Great for ≤70B dense models

16× NVIDIA GB300

$184/hr · FP8

For 1T+ parameter models

runaiicloud

Serverless inference, dedicated GPUs, and training for open models. OpenAI- and Anthropic-compatible APIs.

© 2026 runaii

Platform

Model libraryGPUsPricingCompare providersSavings calculatorDocsServerlessDeploymentsTrainingBatch API

Developers

PlaygroundCookbookCLIAgents / MCPResearch notesUI/UX systemUse casesTutorialsModel advisorBlogCustomersFAQ

Company

EnterpriseStartupsAboutCareersPartnersTrust centerSLAStatusChangelogrunaii chatSupportAPI keysTermsPrivacy