runaiicloud
ModelsGPUsPricingDocsConnectCompareEnterprisePlayground
Log inGet started
runaiicloud docs
status →
⌘⌘K
Get started
Quickstart
API reference
Authentication & API keysStreamingErrorsModels catalog
Billing & limits
Pricing, metering & prompt cachingRate limits & tiers
Bring your own model →Blog ←Get an API key →Changelog
Authentication & API keysPricing, metering & prompt cachingQuickstartRate limits & tiersStreamingErrorsModels catalog

[ docs — api reference ]

runaii.cloud API reference

An OpenAI-compatible, per-token-metered inference API on our own GPU fleet. Swap your base_url, keep your SDK — pay per token with automatic prompt-cache discounts.

Get started
  • Quickstart→
    Make your first metered API call in 60 seconds.
API reference
  • Authentication & API keys→
    Bearer keys, key hygiene, rotation, and revocation.
  • Streaming→
    Server-sent events with a terminal usage chunk — OpenAI SDK compatible.
  • Errors→
    Every error code the API returns and what to do about it.
  • Models catalog→
    Every serverless model, pricing, and context window. Pull it live from the API.
Billing & limits
  • Pricing, metering & prompt caching→
    Per-token billing, cache-hit subsidies, and how to read usage.cost.
  • Rate limits & tiers→
    Current limits, 429 handling, and how caps interact with your balance.
Guides

    new here?

    Quickstart → 60 seconds to your first call

    Make your first metered API call in 60 seconds.

    runaiicloud

    Serverless inference, dedicated GPUs, and training for open models. OpenAI- and Anthropic-compatible APIs.

    © 2026 runaii

    Platform

    Model libraryGPUsPricingCompare providersSavings calculatorDocsServerlessDeploymentsTrainingBatch API

    Developers

    PlaygroundCookbookCLIAgents / MCPResearch notesUI/UX systemUse casesTutorialsModel advisorBlogCustomersFAQ

    Company

    EnterpriseStartupsAboutCareersPartnersTrust centerSLAStatusChangelogrunaii chatSupportAPI keysTermsPrivacy