runaiicloud
ModelsAppsGPUsServerless GPUPricingDocsConnectCompareEnterprisePlayground
Log inGet started

Token-Max · coding plans

A billion tokens a day for your coding agents

Flat monthly plans sized for real agent burn — about $0.02 per 1M tokens on every tier — on GPU cells we own and tune ourselves. Warm-cache fast, OpenAI-compatible, metered live. Capacity is real and finite — so plans are application-gated.

2.1s

788K-token warm resend

3.09s

1M-token warm resend

280–330ms

Warm interactive turns

$0.15/1M in

Flash list price

Solo

$49/mo

100M tokens / day

≈ 3B tokens a month

One heavy agent seat — your daily driver

  • ✓glm-5.3-flash + deepseek-v4-flash
  • ✓15 concurrent requests
  • ✓Daily budget, resets nightly
  • ✓Live meter in the console

Team

$199/mo

300M tokens / day

≈ 9B tokens a month

A small team or parallel agent fleets

  • ✓Everything in Solo
  • ✓50 concurrent requests
  • ✓glm-5.3 flagship at 4× token weight
  • ✓Priority queue lane

Max1B/day

$499/mo

1B tokens / day

≈ 30B tokens a month

Agent fleets and long-horizon runs — a billion-token day

  • ✓Everything in Team
  • ✓100 concurrent requests
  • ✓Capacity-booked: your seat reserves dedicated cell time
  • ✓Warm-cache tuned — agents hit cache, not cold starts

Founding cohort: your rate locks for 12 months. Budgets reset daily (no rollover) and are ceilings, not forecasts — Max seats are capacity-booked: we reserve dedicated cell time for your workload. Need more mid-day? Top up pay-as-you-go credits at list price; plans never block a run.

Why apply, not checkout

We run dedicated cells — H100 and RTX PRO 6000 boxes we kernel-tune ourselves — not a broker's dashboard. Tokens per day is a physical number here, so seats go out in small groups as capacity comes online. Applications keep plan quality high for the builders already on the cells — and your place in the queue is shown the moment you apply.

Apply in two minutes — approved plans start the same week.

Works with your tooling

OpenAI-compatible endpoint today: point Cursor, Cline, Aider, Continue, or Zed at runaii.cloud and keep your workflow. Live token meter in the console, cached reads priced at a fraction of list.

Claude Code speaks Anthropic-format — that endpoint is in active build. Flag it in your application and you're first in line when it lands.

Apply for a plan

Prefer pay-as-you-go? $5 free credits, no application → · Early-stage startup? Get $500 in credits →

runaiicloud

Serverless inference, dedicated GPUs, and training for open models. OpenAI- and Anthropic-compatible APIs.

© 2026 runaii

Platform

Model libraryAppsGPUsServerless GPUPricingToken-Max coding plansCompare providersSavings calculatorDocsServerlessDeploymentsTrainingBatch API

Developers

PlaygroundCookbookCLIAgents / MCPResearch notesUI/UX systemUse casesTutorialsModel advisorBlogCustomersFAQ

Company

EnterpriseStartupsAboutCareersPartnersTrust centerSLAStatusChangelogrunaii chatSupportAPI keysTermsPrivacy