Integrate with Astro

Astro is an AI gateway. You point the LLM client you already have at it, and get

multi-provider access, credential custody and per-turn cost accounting without

changing how you call a model.

Most adoptions are a base URL and a key. If that is all you need, you are three

minutes from done — start at Quickstart and stop reading.


What you get, and what it costs you

You getBecause
One endpoint for every providerAstro speaks both vendor wire formats natively
No vendor keys in your processAstro holds them; you hold an Astro key
Per-turn cost, to the micro-dollarEvery turn writes a ledger row read from the provider's own usage
Budgets that refuse *before* spendingA request that will be refused costs nothing
Failover you never seeA model that fails before the first byte is replaced silently
It costs youMitigation
A network hopAstro adds single-digit milliseconds; your provider call dominates
A dependency on Astro being upYour existing retry/fallback still works — and see Surviving an outage
A key to rotateRotate from the console without a deploy

Quickstart

Already using the Anthropic SDK

import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({
  apiKey: process.env.ASTRO_KEY,          // an Astro key, not an Anthropic one
  baseURL: process.env.ASTRO_BASE_URL,    // e.g. https://astro.example.com
});

// Everything below is unchanged.
const message = await client.messages.create({
  model: "cap:chat.balanced",
  max_tokens: 512,
  messages: [{ role: "user", content: "Summarise this in one line." }],
});

Already using the OpenAI SDK

from openai import OpenAI

client = OpenAI(
    api_key=os.environ["ASTRO_KEY"],
    base_url=os.environ["ASTRO_BASE_URL"] + "/v1",
)

completion = client.chat.completions.create(
    model="cap:chat.balanced",
    messages=[{"role": "user", "content": "Summarise this in one line."}],
)

curl

curl "$ASTRO_BASE_URL/v1/messages" \
  -H "x-api-key: $ASTRO_KEY" \
  -H "content-type: application/json" \
  -d '{
    "model": "cap:chat.balanced",
    "max_tokens": 128,
    "messages": [{ "role": "user", "content": "hello" }]
  }'
Rolling back is unsetting one variable. Write your client so the base URL is
process.env.ASTRO_BASE_URL ?? undefined — absent falls through to the vendor
default with no deploy.

Ask for a capability, not a model

cap:chat.balanced rather than claude-haiku-4-5.

A capability is a promise about *what the model must be able to do* — context

window, tool use, vision, residency — resolved to a concrete model at request time.

It survives a model being retired, and it lets an operator move you onto something

cheaper or faster without you shipping anything.

GET /v1/models lists both. Pin a model id when you genuinely need that model; use a

capability the rest of the time.


What comes back

Every response carries headers worth logging:

HeaderUse
x-astro-request-idQuote this in a support request. It is the join to the ledger row, the audit entry and the logs
x-astro-served-modelWhat actually answered — may differ from what you asked for
x-astro-cost-microsInteger micro-USD for this turn

Errors

Astro returns your vendor's error envelope, so your SDK raises the typed exception

it already raises. The status codes that mean something specific:

StatusMeaningRetry?
401No key, or a key this deployment does not knowNo
402A budget refused this turn. It never reached a provider and cost nothingNo — it will refuse again
429Rate limitedYes, with backoff
502The provider failedYes, if your workload tolerates it
503No credential is wired for a model that could otherwise serveNo — tell an operator

A 402 carries the numbers to act on:

{
  "type": "error",
  "error": { "type": "billing_error", "message": "budget exceeded" },
  "astro": {
    "limit_micros": 25000000,
    "committed_micros": 24999100,
    "shortfall_micros": 4200
  }
}

shortfall_micros is the number that matters. Large means the ceiling is wrong for

your workload; small means a nearly-full budget met an ordinary request and the

window will reset.


Streaming

Set stream: true. You get your vendor's SSE format — Anthropic's named events on

/v1/messages, chat.completion.chunk frames terminated by the literal

data: [DONE] on /v1/chat/completions.

Two behaviours worth knowing:

like any other failed request.

closes.** Astro does not fail over mid-stream: a second model would continue

someone else's sentence.


Surviving an outage

Astro is a dependency. Plan for it being down the way you would any other.

If you already have fallback logic, keep it. The safest adopters put Astro *last*

in an existing fallback chain first, then promote it once it has proven itself.

If you want offline-capable tests, use @astro/sdk, which short-circuits inside

the process with no key and no network:

import { createClient } from "@astro/sdk";

const astro = createClient();   // mode: "auto" — live if credentials resolve, offline otherwise

const res = await astro.chatSoft({
  model: "cap:chat.balanced",
  messages: [{ role: "user", content: "hello" }],
});

if (!res.ok) {
  // Never throws — including on programmer error. Branch on a kind, not a message.
  if (res.error.kind === "over_budget") tellTheUser();
  else if (res.error.retryable) scheduleRetry();
}

Offline responses are deterministic — the same request gives the same answer on

every machine — so your suite can assert on them. Every response carries source

(live · offline · replay) so you can always tell what answered.

Set ASTRO_STRICT=1 in production. It makes an accidental fall-back to offline
fatal instead of silent. A service quietly answering from a stub because a key was
missing in staging looks healthy while being entirely fictional.

Record real answers once and replay them forever:

import { createClient, fileCassettes } from "@astro/sdk";

const cassettes = fileCassettes("./test/cassettes");
const recording = createClient({ mode: "record", cassettes });   // calls once, saves
const replaying = createClient({ mode: "replay", cassettes });   // never calls

A replay miss is an error, never a live call — otherwise a suite that was meant

to be hermetic quietly starts burning keys.


Rate limits, budgets and cost

Your key is scoped to a project. That project may have a budget with a window

(turn · hour · day · month) and an action when exceeded:

Ask your operator which applies. If you are seeing 402s, the

refusals screen in their console shows exactly how far over each request went.


Checklist before you ship


API reference

The full public surface — every route, every field, with examples — is in the

API reference.