runaiicloud
ModelsGPUsPricingDocsConnectCompareEnterprisePlayground
Log inGet started
← Model library

Kimi K3

NEW

Frontier open-weights model for coding, reasoning, and long context.

Try in Playground

Input / 1M

$3.00

Output / 1M

$15.00

Context

1M

VisionmoonshotaiLoRA tunableServerless

Benchmarks

Published evaluation scores (higher is better). Measured output speed on our fleet, Standard tier.

SWE-bench Verified

71.8%

Real GitHub issues

GPQA Diamond

75.7%

PhD-level science

MMLU Pro

82.1%

Knowledge & reasoning

AIME 2025

90.2%

Competition math

IFEval

89.5%

Instruction following

Output speed

96

tok/s on runaii

Ways to serve this model

ServerlessPer-token, zero cold starts. From $3.00/M input.Start serverless →Dedicated deploymentReserved GPUs from $40/hr. Full speed control, scale-to-zero, your region.Create deployment →LoRA fine-tuneTrain on your data from $/1M tokens; checkpoints deploy in seconds.Start training →

Live model traffic — last 24h

Requests to runaii/kimi-k3 across the fleet, from the live meter.

Requests

…

Tokens

…

Spend

…

Quickstart

OpenAI SDK (change the base URL) or the Anthropic SDK — both speak runaii.cloud natively.

curl https://api.runaii.cloud/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $RUNAII_API_KEY" \
  -d '{
    "model": "runaii/kimi-k3",
    "messages": [{"role": "user", "content": "Say hello in Spanish"}]
  }'
liveMaking requestMaking request
runaiicloud

Serverless inference, dedicated GPUs, and training for open models. OpenAI- and Anthropic-compatible APIs.

© 2026 runaii

Platform

Model libraryGPUsPricingCompare providersSavings calculatorDocsServerlessDeploymentsTrainingBatch API

Developers

PlaygroundCookbookCLIAgents / MCPResearch notesUI/UX systemUse casesTutorialsModel advisorBlogCustomersFAQ

Company

EnterpriseStartupsAboutCareersPartnersTrust centerSLAStatusChangelogrunaii chatSupportAPI keysTermsPrivacy