runaiicloud
ModelsGPUsPricingDocsConnectCompareEnterprisePlayground
Log inGet started
← All GPUs
DIEHBMINTERPOSERNVSWITCHSUBSTRATESUBSTRATE · POWER & THERMALINTERPOSER · NVLink-C2C FABRICHBMHBMHBMHBMB200NVFP4 · 4×NVSWITCH

Efficient · NVFP4

4× NVIDIA B200 180GB

$40/hr $0.0111/sec · $29,200/mo at 24/7

Great for ≤70B dense models. ≤70B dense models, evals.

VRAM pool

180GB ×4 = 720GB

Position

vs L40S 4×: ~2.5× inference speed

Deploy this →Configure sizeCompare providers →

What you're renting

Compute die

NVIDIA B200 180GB — NVFP4 tensor cores, tuned for ≤70b dense models, evals.

HBM stack

180GB ×4 = 720GB of high-bandwidth attached memory — the whole model fits, no offload.

Interposer / NVLink

NVLink-C2C fabric so the 4 GPUs act like one pool.

Per-second metering

$40/hr → 0.0111/sec. Pause the deployment and the meter hits $0.

Scale-to-zero

Autoscale + region pins (GLOBAL / US / EU / APAC) on the control plane.

Reserve it in minutesDeploy →

Other presets

8× NVIDIA B300 288GB

$96/hr · FP8

Max quality for frontier dense models

8× NVIDIA B200 180GB

$80/hr · NVFP4

Best tokens/$ for large MoE models

4× NVIDIA B300 288GB

$48/hr · NVFP4

Smallest frontier-capable footprint

16× NVIDIA GB300

$184/hr · FP8

For 1T+ parameter models

runaiicloud

Serverless inference, dedicated GPUs, and training for open models. OpenAI- and Anthropic-compatible APIs.

© 2026 runaii

Platform

Model libraryGPUsPricingCompare providersSavings calculatorDocsServerlessDeploymentsTrainingBatch API

Developers

PlaygroundCookbookCLIAgents / MCPResearch notesUI/UX systemUse casesTutorialsModel advisorBlogCustomersFAQ

Company

EnterpriseStartupsAboutCareersPartnersTrust centerSLAStatusChangelogrunaii chatSupportAPI keysTermsPrivacy