[ public beta ] — $5 free credits on every new account
runaiicloud
ModelsAppsGPUsPricingDocsConnectCompareEnterprisePlayground
Log inGet started

[ runaii cloud — serverless · dedicated · training ]

Own your
inference.

Serverless tokens, dedicated Blackwell GPUs, and training — one API your SDK already speaks. Cached prefixes up to 97% off.

Start building — $5 freeRead the docs

no card · no sales call · how we benchmarked the incumbents

runaii — zshexit 0 · 312ms
curl https://api.runaii.cloud/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $RUNAII_API_KEY" \
  -d '{
    "model": "runaii/glm-5.3",
    "messages": [{"role": "user", "content": "Say hello in Spanish"}]
  }'
liveMaking requestMaking request
OpenAI + Anthropic compatibletry the live playground →
<0ms
P50 cold start, H100
0 tok/s
peak MoE throughput
0
models in the catalog
$0
free, no card
$0.000/M
cached input · up to 97% off

gateway live·billed to the micro·status →

[ featured — deepseek-v4-flash ]

Flash-class speed. $0.02 cached.

The high-volume workload we actually serve — 131K context, 218 tok/s measured on our fleet, and cache-hit input under three cents. One key, every model.

Try it — $5 freeModel card →

input

$0.22

/ 1M tokens

cached

$0.022

/ 1M tokens

output

$0.66

/ 1M tokens

B — fleet speeds

Tokens per second, on our fleet

Median output speed, Standard tier. Same architecture, honest numbers.

GLM 5.3 Flash412 tok/s66K ctx$0.15/1M · ⚡0.030 cached
DeepSeek V4.1 Flash218 tok/s131K ctx$0.22/1M · ⚡0.022 cached
GLM-5.3118 tok/s203K ctx$1.40/1M · ⚡0.260 cached
Qwen3.8-Max104 tok/s33K ctx$2.00/1M

how we benchmarked the incumbents →

[ open models — one api, one key ]

all 5 →
GLGLM-5.3$1.40/1M· ⚡$0.260 cachedGFGLM 5.3 Flash$0.15/1M· ⚡$0.030 cachedDeepSeek V4.1 Flash$0.22/1M· ⚡$0.022 cachedQMQwen3.8-Max$2.00/1MQFQwen3.8-27B Flash$0.40/1MGLGLM-5.3$1.40/1M· ⚡$0.260 cachedGFGLM 5.3 Flash$0.15/1M· ⚡$0.030 cachedDeepSeek V4.1 Flash$0.22/1M· ⚡$0.022 cachedQMQwen3.8-Max$2.00/1MQFQwen3.8-27B Flash$0.40/1M
runaii/glm-5.3 $1.40/$4.40/1M · $0.26 cached/runaii/deepseek-v4-flash $0.22/$0.66/1M · $0.022 cached/runaii/qwen3.8-flash $0.40/$1.60/1M/runaii/glm-5.3 $1.40/$4.40/1M · $0.26 cached/runaii/deepseek-v4-flash $0.22/$0.66/1M · $0.022 cached/runaii/qwen3.8-flash $0.40/$1.60/1M/
A — 01/03

The stack

Three primitives. One control plane. No replatforming between stages.

01 / Serverless

Per-token inference. Zero capacity planning.

Fifteen open models behind one OpenAI-compatible URL; cached prefixes bill for almost nothing.

1M ctx · streaming · tools →

02 / Deployments

Your weights on B200s, billed per second.

Reserve Blackwell-class GPUs with autoscaling and scale-to-zero. Your VPC, your region pin, your control plane.

B200 · B300 · GB300 →

03 / Training

LoRA to RL, then straight to prod.

Managed SFT/DPO per 1M tokens, dedicated RL per GPU-second. Every checkpoint deploys to inference in seconds.

SFT · DPO · RL →

[ metered in real time ]

Watch every token land.

Per-request receipts: tokens in/out, cache hit rate, latency, and the exact micros charged — streamed to your console the moment a call finishes.

Open the consolereceipts doc →
runaii — live usagestreaming

req/min

0

tok/s

0

p50

0ms

14:20 200 glm-5.3 1.4k tok cached 62% $0.00041 182ms

15:23 200 deepseek-v4-flash 8.1k tok stream $0.01180 1.1s

16:26 200 glm-5.3 512 tok stream $0.00033 96ms

17:29 429 rate-limited — retry 60s $0.00000 —

18:32 200 qwen3.8-flash 2.2k tok cached 31% $0.00290 640ms

[ try it — watch a request happen ]

Type a prompt. See the bill.

A working preview of the request flow — route, stream, meter. The same path your code takes, minus the API key.

POST /v1/chat/completions · deepseek-v4-flashdemo
$ Explain vector search in one line
— hit Run to watch tokens stream —
Making request—

Visual demo of the request flow — route, stream, meter. Run it against your own key in the playground →

[ gpu cloud — runaii, not a reseller ]

Silicon, priced per second.

One platform. Blackwell managed presets for serving, or raw GPU VMs (T4 · L4 · RTX PRO 6000 · A100 · H100) with Spot discounts and Standard/Premium network egress. Launch in one click, bill by the minute.

8× NVIDIA B300 288GBFP8Max quality for frontier dense models$96/hr $0.0267/s
8× NVIDIA B200 180GBNVFP4Best tokens/$ for large MoE models$80/hr $0.0222/s
4× NVIDIA B300 288GBNVFP4Smallest frontier-capable footprint$48/hr $0.0133/s
4× NVIDIA B200 180GBNVFP4Great for ≤70B dense models$40/hr $0.0111/s
16× NVIDIA GB300FP8For 1T+ parameter models$184/hr $0.0511/s

1× T4

8vCPU · 30G

Cheap inference, transcoding, and CUDA experiments

$0.44/hrspot $0.29
Launch →

1× L4popular

8vCPU · 32G

Best price/perf for LLM inference up to ~13B and video

$0.87/hrspot $0.57
Launch →

4× L4

24vCPU · 96G

Multi-worker inference farms on one VM

$3.60/hrspot $2.34
Launch →

1× RTX PRO 6000

24vCPU · 96G

96GB Blackwell workstation-class VRAM, single GPU

$2.40/hrspot $1.56
Launch →

1× A100

12vCPU · 85G

Ampere workhorse for training and dense inference

$2.93/hrspot $1.90
Launch →

8× A100

96vCPU · 680G

Full-node A100 training

$23.46/hrspot $15.25
Launch →
Browse all GPUsLaunch a GPU VM →
B — receipts

Us vs everybody

full 7-way table →

Serverless APIs

15 open models, one URL

Modal: SDK-first · Baseten: curated set

Dedicated GPUs

B200/B300/GB300, per-second

RunPod: 30+ SKUs · Koyeb: A100/L40S

Training loop

Checkpoint → endpoint in seconds

Fireworks: the same loop, closed walls

Start cost

$5 free, no card, no sales call

Baseten: sales-led · Koyeb: credits via program

Model library

all 5 →
GLM-5.3NEWFrontier agentic model. Top-tier coding, tool use, and long-context reas…$1.40/$4.40⚡ $0.260/M cached203K ctxGLM 5.3 FlashNEWUltra-fast multimodal workhorse for chat, extraction, and routing.…$0.15/$0.50⚡ $0.030/M cached66K ctxDeepSeek V4.1 FlashServed on DeepSeek V4.1-Flash weights — speed-tuned for high-volume chat…$0.22/$0.66⚡ $0.022/M cached131K ctxQwen3.8-MaxQwen flagship. Multilingual strength and dense world knowledge.…$2.00/$6.0033K ctxQwen3.8-27B FlashBalanced model for production chat — served on Qwen3.8-27B weights.…$0.40/$1.6033K ctx

Ship tonight

cookbook →

example_01.py

Streaming chat in 9 lines

OpenAI SDK, new base URL, token-by-token SSE.

model="runaii/deepseek-v4-flash"

example_02.py

RAG that cites sources

Embed → retrieve → rerank → answer, cached.

model="runaii/qwen3-embedding-8b"

example_03.py

Tool-calling agent loop

Declare tools, loop 8 deep, cap the spend.

model="runaii/glm-5.3"

[ works with — your stack, unchanged ]

OpenAI SDKAnthropic SDKLangChainVercel AI SDKLlamaIndexCursorClaude Codecurl

what early-access beta partners tell us

-60%

model bill after switching

We moved our main agent to an open model and nobody noticed — except the bill dropped 60%.
platform team · beta partner

<1s

cold start on H100s

Cold starts under a second on H100s. Evals went from overnight to over lunch.
ML lead · inference beta

1 day

base-URL migration to prod

One base URL change and we were compatible. The playground sold the team in a day.
founding engineer · early access
customer stories →

Start with $5.
Stay for the bill.

Get started →Pricing
runaiicloud

Serverless inference, dedicated GPUs, and training for open models. OpenAI- and Anthropic-compatible APIs.

© 2026 runaii

Platform

Model libraryAppsGPUsPricingCompare providersSavings calculatorDocsServerlessDeploymentsTrainingBatch API

Developers

PlaygroundCookbookCLIAgents / MCPResearch notesUI/UX systemUse casesTutorialsModel advisorBlogCustomersFAQ

Company

EnterpriseStartupsAboutCareersPartnersTrust centerSLAStatusChangelogrunaii chatSupportAPI keysTermsPrivacy