runaiicloud
ModelsGPUsPricingDocsConnectCompareEnterprisePlayground
Log inGet started

Model library

Every open weight.
One API.

Try in playground →

15

serverless models

1M

max context window

B200 · B300

Blackwell GPUs on demand

$1

free credits, no card

Start here

editor's picks — real prices, real context
#1NEW

GLM-5.3

z.ai · LLM

Frontier agentic model. Top-tier coding, tool use, and long-context reasoning.

$1.40/$4.401M ctx
#2NEW

Kimi K3

moonshotai · Vision

Frontier open-weights model for coding, reasoning, and long context.

$3.00/$15.001M ctx
#3NEW

GLM 5.3 Flash

z.ai · Vision

Ultra-fast multimodal workhorse for chat, extraction, and routing.

$0.15/$0.501M ctx

⚡ Make a request — live builder

Pick a model, tune the knobs, copy the code or simulate a streaming reply.

Open this in playground →
Temperature · balanced0.7
Top-p · nucleus sampling0.90
Max tokens · 1024 tokens1024

Simulated streaming reply

Hit “Run demo stream” to preview token-by-token output…

request.py
from openai import OpenAI

client = OpenAI(
    api_key="$RUNAII_API_KEY",
    base_url="https://api.runaii.cloud/v1",
)

resp = client.chat.completions.create(
    model="runaii/glm-5.3",
    temperature=0.7,
    top_p=0.9,
    max_tokens=1024,
    messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)

~320ms

TTFT (typical)

211 tok/s

Peak throughput

1M

Context

📡 Console activity — demo preview

simulated

Why builders switch

  • ✓ One base URL for 15+ open models
  • ✓ Cached prefixes billed up to 97% off
  • ✓ Batch API at 50% off, same gateway
  • ✓ $1 free credits — no card required
Compare pricingRead the docs

All models

GLM-5.3

z.ai

NEW

Frontier agentic model. Top-tier coding, tool use, and long-context reasoning.

$1.40 / $4.401M ctx
chat · tools · streaming
◆ LLMFn callingTunableCached $0.26
Try in playgroundDetails →

GLM 5.3 Flash

z.ai

NEW

Ultra-fast multimodal workhorse for chat, extraction, and routing.

$0.15 / $0.501M ctx
vision
◉ VisionFn callingCached $0.03
Try in playgroundDetails →

Kimi K3

moonshotai

NEW

Frontier open-weights model for coding, reasoning, and long context.

$3.00 / $15.001M ctx
vision
◉ VisionTunable
Try in playgroundDetails →

DeepSeek V4 Pro

deepseek

SWE-bench leader. Strong math, code, and agentic tool use.

$1.32 / $3.961M ctx
chat · tools · streaming
◆ LLMFn callingTunableCached $0.04
Try in playgroundDetails →

DeepSeek V4 Flash

deepseek

Speed-tuned V4 for high-volume pipelines.

$0.22 / $0.661M ctx
chat · tools · streaming
◆ LLMCached $0.02
Try in playgroundDetails →

Qwen3.8-Max

qwen

Qwen flagship. Multilingual strength and dense world knowledge.

$2.00 / $6.00262K ctx
chat · tools · streaming
◆ LLMFn calling
Try in playgroundDetails →

Qwen3.8 Flash

qwen

Balanced multimodal model for production chat.

$0.40 / $1.60262K ctx
vision
◉ Vision
Try in playgroundDetails →

OpenAI gpt-oss-120b

openai

OpenAI's open-weights 120B MoE. Efficient reasoning per dollar.

$0.15 / $0.60131K ctx
chat · tools · streaming
◆ LLMFn calling
Try in playgroundDetails →

OpenAI gpt-oss-20b

openai

Small open-weights model for classification and routing.

$0.05 / $0.20131K ctx
chat · tools · streaming
◆ LLM
Try in playgroundDetails →

MiniMax M3

minimax

Long-context generalist with aggressive pricing.

$0.30 / $1.20512K ctx
chat · tools · streaming
◆ LLM
Try in playgroundDetails →

Llama 3.3 70B Instruct

meta

The dependable classic. Broad ecosystem support.

$0.23 / $0.85131K ctx
chat · tools · streaming
◆ LLMFn callingTunable
Try in playgroundDetails →

Qwen3 Embedding 8B

qwen

State-of-the-art multilingual embeddings.

$0.10 / M41K ctx
embedding
⌗ Embedding
Try in playgroundDetails →

Qwen3 Reranker 8B

qwen

Cross-encoder reranking for RAG pipelines.

$0.20 / M41K ctx
rerank
⇅ Rerank
Try in playgroundDetails →

Whisper V3 Large

openai

Speech-to-text across 99 languages.

$0.25 / M
audio
◍ Audio
Try in playgroundDetails →

FLUX.1 Kontext Pro

blackforest

Context-aware image generation and editing.

$0.04 / M
image
▤ Image
Try in playgroundDetails →

Not sure which model? Answer 2 questions.

The advisor picks 3 open models sized to your use case, cost, and context needs — every price on this page.

Step 1 / 2

What are you building?

Found your model? Try it free.

Open the playground with any model preloaded — $1 credit covers thousands of calls.

Get started →Pricing
runaiicloud

Serverless inference, dedicated GPUs, and training for open models. OpenAI- and Anthropic-compatible APIs.

© 2026 runaii

Platform

Model libraryGPUsPricingCompare providersSavings calculatorDocsServerlessDeploymentsTrainingBatch API

Developers

PlaygroundCookbookCLIAgents / MCPResearch notesUI/UX systemUse casesTutorialsModel advisorBlogCustomersFAQ

Company

EnterpriseStartupsAboutCareersPartnersTrust centerSLAStatusChangelogrunaii chatSupportAPI keysTermsPrivacy