runaiicloud
ModelsAppsGPUsServerless GPUPricingDocsConnectCompareEnterprisePlayground
Log inGet started
← All GPUs
DIEHBMINTERPOSERNVSWITCHSUBSTRATESUBSTRATE · POWER & THERMALINTERPOSER · NVLink-C2C FABRICHBMHBMHBMHBML4FP8 · 8×NVSWITCH

Efficient ×8 · FP8

8× NVIDIA L4 24GB

$7.2/hr $0.0020/sec · $5,256/mo at 24/7

g2-standard-96 — full-node L4. g2-standard-96 — full-node L4.

VRAM pool

8× NVIDIA L4 24GB

Position

Priced per GPU-hour — see the launch wizard for live rates

Deploy this →Configure sizeCompare providers →

What you're renting

Compute die

NVIDIA L4 24GB — FP8 tensor cores, tuned for g2-standard-96 — full-node l4.

HBM stack

8× NVIDIA L4 24GB of high-bandwidth attached memory — the whole model fits, no offload.

Interposer / NVLink

NVLink-C2C fabric so the 8 GPUs act like one pool.

Per-second metering

$7.2/hr → 0.0020/sec. Pause the deployment and the meter hits $0.

Scale-to-zero

Autoscale + region pins (GLOBAL / US / EU / APAC) on the control plane.

Reserve it in minutesDeploy →

Other presets

1× NVIDIA T4 16GB

$0.44/hr · FP16

Cheapest CUDA — experiments, small models

1× NVIDIA L4 24GB

$0.87/hr · FP8

Best price/perf up to ~13B + video

2× NVIDIA L4 24GB

$1.8/hr · FP8

g2-standard-24 — two-worker inference

4× NVIDIA L4 24GB

$3.6/hr · FP8

g2-standard-48 — multi-worker farm on one VM

1× NVIDIA RTX PRO 6000 Blackwell (1/8 GPU · 12GB)

$0.3/hr · FP8

g4-standard-6 — Blackwell slice for light jobs

1× NVIDIA RTX PRO 6000 Blackwell (1/4 GPU · 24GB)

$0.6/hr · FP8

g4-standard-12 — quarter GPU

1× NVIDIA RTX PRO 6000 Blackwell (1/2 GPU · 48GB)

$1.2/hr · FP8

g4-standard-24 — half GPU

1× NVIDIA RTX PRO 6000 Blackwell 96GB

$2.4/hr · FP8

g4-standard-48 — try any app from $2.40/hr (park/wake receipted D27)

2× NVIDIA RTX PRO 6000 Blackwell 96GB

$4.8/hr · FP8

g4-standard-96 — two Blackwell GPUs

4× NVIDIA RTX PRO 6000 Blackwell 96GB

$9.6/hr · FP8

g4-standard-192 — four GPUs

8× NVIDIA RTX PRO 6000 Blackwell 96GB

$19.2/hr · FP8

g4-standard-384 — full-node Blackwell

8× NVIDIA B300 288GB

$96/hr · FP8

Max quality for frontier dense models

8× NVIDIA B200 180GB

$80/hr · NVFP4

Best tokens/$ for large MoE models

4× NVIDIA B300 288GB

$48/hr · NVFP4

Smallest frontier-capable footprint

4× NVIDIA B200 180GB

$40/hr · NVFP4

Great for ≤70B dense models

16× NVIDIA GB300

$184/hr · FP8

For 1T+ parameter models

runaiicloud

Serverless inference, dedicated GPUs, and training for open models. OpenAI- and Anthropic-compatible APIs.

© 2026 runaii

Platform

Model libraryAppsGPUsServerless GPUPricingToken-Max coding plansCompare providersSavings calculatorDocsServerlessDeploymentsTrainingBatch API

Developers

PlaygroundCookbookCLIAgents / MCPResearch notesUI/UX systemUse casesTutorialsModel advisorBlogCustomersFAQ

Company

EnterpriseStartupsAboutCareersPartnersTrust centerSLAStatusChangelogrunaii chatSupportAPI keysTermsPrivacy