[ runaii cloud — serverless · dedicated · training ]
Serverless tokens, dedicated Blackwell GPUs, and training — one API your SDK already speaks. Cached prefixes up to 97% off.
no card · no sales call · how we benchmarked the incumbents
curl https://api.runaii.cloud/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $RUNAII_API_KEY" \
-d '{
"model": "runaii/glm-5.3",
"messages": [{"role": "user", "content": "Say hello in Spanish"}]
}'gateway livebilled to the microstatus →
[ featured — glm-5.3-flash ]
The flash-class workload everyone benchmarks — we ship it with 1M context, 412 tok/s on our fleet, and cache-hit input at under a nickel. One key, every model.
input
$0.15
/ 1M tokens
cached
$0.030
/ 1M tokens
output
$0.50
/ 1M tokens
| GLM 5.3 Flash | 412 tok/s | 1M ctx | $0.15/1M · ⚡0.030 cached |
| DeepSeek V4 Flash | 385 tok/s | 1M ctx | $0.22/1M · ⚡0.022 cached |
| OpenAI gpt-oss-120b | 260 tok/s | 131K ctx | $0.15/1M |
| MiniMax M3 | 190 tok/s | 512K ctx | $0.30/1M |
| Llama 3.3 70B Instruct | 175 tok/s | 131K ctx | $0.23/1M |
| GLM-5.3 | 118 tok/s | 1M ctx | $1.40/1M · ⚡0.260 cached |
| Qwen3.8-Max | 104 tok/s | 262K ctx | $2.00/1M |
| Kimi K3 | 96 tok/s | 1M ctx | $3.00/1M |
[ open models — one api, one key ]
all 15 →[ metered in real time ]
Per-request receipts: tokens in/out, cache hit rate, latency, and the exact micros charged — streamed to your console the moment a call finishes.
[ try it — watch a request happen ]
A working preview of the request flow — route, stream, meter. The same path your code takes, minus the API key.
$ Explain vector search in one line— hit Run to watch tokens stream —Visual demo of the request flow — route, stream, meter. Run it against your own key in the playground →
| 8× NVIDIA B300 288GB | FP8 | Max quality for frontier dense models | $96/hr $0.0267/s |
| 8× NVIDIA B200 180GB | NVFP4 | Best tokens/$ for large MoE models | $80/hr $0.0222/s |
| 4× NVIDIA B300 288GB | NVFP4 | Smallest frontier-capable footprint | $48/hr $0.0133/s |
| 4× NVIDIA B200 180GB | NVFP4 | Great for ≤70B dense models | $40/hr $0.0111/s |
| 16× NVIDIA GB300 | FP8 | For 1T+ parameter models | $184/hr $0.0511/s |
Serverless APIs
15+ open models, one URL
Modal: SDK-first · Baseten: curated set
Dedicated GPUs
B200/B300/GB300, per-second
RunPod: 30+ SKUs · Koyeb: A100/L40S
Training loop
Checkpoint → endpoint in seconds
Fireworks: the same loop, closed walls
Start cost
$1 free, no card, no sales call
Baseten: sales-led · Koyeb: credits via program
[ works with — your stack, unchanged ]
-60%
model bill after switching
We moved our main agent to an open model and nobody noticed — except the bill dropped 60%.
<1s
cold start on B200s
Cold starts under a second on B200s. Evals went from overnight to over lunch.
1 day
base-URL migration to prod
One base URL change and we were compatible. The playground sold the team in a day.