Docs
Status
OverviewQuickstartSetup promptsThe core loopAuthenticationOverviewModelsThe waterfallAdding modelsAnthropic APIErrorsIntegrate the gatewayCost APIAccount APICoding agentsCredits & billingTelemetryAPI reference

Get started

  • Overview
  • Quickstart
  • Setup prompts
  • The core loop
  • Authentication

Guides

  • Overview
  • Models
  • The waterfall
  • Adding models
  • Anthropic API
  • Errors

Integrations

  • Integrate the gateway
  • Cost API
  • Account API
  • Coding agents

Billing & usage

  • Credits & billing
  • Telemetry

Reference

  • API reference
PreviousCoding agentsNextTelemetry

Billing & usage

Credits & billing

Every model is paid for through one of two lanes, and the gateway adds no markup on either. Platform-funded calls draw down your credits; bring-your-own-key calls are billed by the provider directly.

Two lanes

  • Platform-funded: our credits, priced from the public catalog. Each call draws down your balance and is metered as cost_micro_usd.
  • Pass-through (BYOK): your own provider key. The provider bills you directly, so these calls never draw credits; they are metered as estimated_cost_micro_usd for attribution only.

Which lane a model rides is decided per-provider by its waterfall: a deployment backed by one of your provider connections is pass-through; a platform-seeded deployment is platform-funded. Either way, zero markup.

Credits

A credit is the platform’s spendable unit, pegged at a flat one cent. It is what everything is priced in, so display, checkout, and spend never drift between dollars and tokens. Today a credit is simply a cent under a friendlier name: the only thing that draws credits is routed token usage, which stays zero margin (platform-funded calls draw credits at the provider’s catalog price, with nothing added on top).

You get credits two ways, both at one cent each. The Free plan refreshes a monthly allotment. A Pro plan grants a larger monthly allotment and unlocks the Pro features. One-off top-ups buy credits without a plan, at the same flat rate. There is no markup on routed tokens; any margin comes from plans, not from a credit spread.

A new organization starts with a welcome credit grant. Your balance is the credit granted minus your billable (platform-funded) spend; pass-through usage does not count against it. Balance, spend, adding credits, and auto-recharge live in the dashboard at Credits.

Balance and top-ups are dashboard (web-session) actions, not API-key actions. An agent tracks its own consumption through usage instead: read GET /api/gateway/usage/daily for spend by day, model, or member. See Telemetry.

Provider accounts and rotation

You can connect more than one account for the same provider — two Anthropic keys, two OpenAI organizations — each under its own handle. They form a pool: the gateway serves your traffic on the first account in your order, and rotates to the next one when an account runs out of quota, or is rate-limited in a sustained way (a burst of throttles over fifteen minutes; a single throttle never rotates, because switching accounts busts the prompt cache you have built on the current one). Rotation is a verdict written from your own traffic every five minutes; a later successful key check re-admits the account.

Manage the pool on the Credits page: drag accounts to set the order, switch each account’s “rotate when out of quota” and “rotate on sustained rate limit” off to fail on it instead of spending on a sibling, and read every account’s usage on its own key. The same controls are one call for an agent holding your org key: GET /api/orgs/{org_id}/provider-connections/accounts/usage, POST /api/orgs/{org_id}/provider-connections/{provider}/accounts/{setup_alias}/routing, and POST /api/orgs/{org_id}/provider-connections/reorder.

Accounts are provider credentials you own. Spend on them is the provider’s bill, shown here for visibility; it never draws platform credits.

Spend controls

Spend is bounded at three levels, all configured in the dashboard:

  • Per-key limits: a daily platform-funded spend cap, a requests-per-minute ceiling, and a tokens-per-minute ceiling on each API key.
  • Budgets: a spend ceiling scoped to the whole team, an API key, a model, an identity, or a routing pool, for a given month or as a recurring cap.
  • Spend alerts: an email when monthly org spend or a budget crosses a threshold.

A key can read its own effective limits over the API. GET /api/gateway/keys/<api_key_id>/limits returns the three ceilings with platform defaults folded in; a null value means uncapped, and source is explicit when set on the key or default otherwise. Setting limits is an admin dashboard action.

GET /api/gateway/keys/{api_key_id}/limits
curl "https://api-pr-1743.preview.experientiallabs.ai/api/gateway/keys/$API_KEY_ID/limits" \
-H "Authorization: Bearer $EXPLABS_API_KEY"
fieldMeaning
daily_spend_cap_micro_usdMax platform-funded spend per day for this key (micro-USD).
requests_per_minuteRequest-rate ceiling for this key.
tokens_per_minuteToken-rate (TPM) ceiling for this key.

Free tiers and credits overflow

Some platform-funded models carry a promotional free daily tier (today gpt-6-astra and claude-fable-5.1); the model page shows the tier as its own "Free tier" rung above the regular pay-as-you-go rate. Eligibility is a saved card and one settled $1 charge on the organization — adding a card alone is not a charge. Each tier has per-org daily and hourly token allowances (the model page names the exact numbers); cached input tokens do not count against them.

  • Past the allowance the request answers 429 insufficient_quota with a free_limit_reached message and does not spend credits. The daily allowance resets at 00:00 UTC, the hourly one at the top of the hour.
  • Credits overflow changes that: when on, requests past the free limit bill the overage to your credits at list price instead of throttling. The switch is org-wide (it covers every free model). It is off for an org that has never verified a card and turns on automatically once your org has a card on file and the settled $1 verification, or at your first real payment — a Pro subscription, a credit top-up, or an auto-recharge — which the dashboard shows you once, right after checkout. If you turned it off, you can turn it back on by hand once your org clears that same bar (the one the free tiers' card gate reads; a purchase or Pro satisfies it): an org admin flips it on any free model's page ("Past the free limit" → Use credits), or an agent can POST /api/credits-overflow on the web host with Authorization: Bearer xpl_... (enable-only, idempotent). Before that, both answer 402 verification_required — add a card and complete the $1 verification to unlock it. Turn it off again from the same model-page row.
  • service_tier: "flex" on a Chat Completions or Responses request to gpt-5.6-solforwards OpenAI's flex tier and bills its rate (50% of base) at cost; on a model without tier pricing it answers 400 unsupported_capability.

When you run out

When your credit balance, a spend limit, or a free tier is exhausted, calls fail with 429 insufficient_quota; the message says which: key_daily_cap, a budget, insufficient_credits, free_limit_reached, free_tier_requires_payment (add a card and a $1 charge), or promo_byok_only (the free tier is spent and your balance cannot cover the request). It is not transient: retrying does not clear it.

How to recover

  • Add credits or raise a limit in the dashboard (platform-funded lane).
  • Or move the model to the pass-through lane by connecting a provider key.

The full error contract is in Errors.