Integrate with Astro
Astro is an AI gateway. You point the LLM client you already have at it, and get
multi-provider access, credential custody and per-turn cost accounting without
changing how you call a model.
Most adoptions are a base URL and a key. If that is all you need, you are three
minutes from done — start at Quickstart and stop reading.
What you get, and what it costs you
| You get | Because |
|---|---|
| One endpoint for every provider | Astro speaks both vendor wire formats natively |
| No vendor keys in your process | Astro holds them; you hold an Astro key |
| Per-turn cost, to the micro-dollar | Every turn writes a ledger row read from the provider's own usage |
| Budgets that refuse *before* spending | A request that will be refused costs nothing |
| Failover you never see | A model that fails before the first byte is replaced silently |
| It costs you | Mitigation |
|---|---|
| A network hop | Astro adds single-digit milliseconds; your provider call dominates |
| A dependency on Astro being up | Your existing retry/fallback still works — and see Surviving an outage |
| A key to rotate | Rotate from the console without a deploy |
Quickstart
Already using the Anthropic SDK
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
apiKey: process.env.ASTRO_KEY, // an Astro key, not an Anthropic one
baseURL: process.env.ASTRO_BASE_URL, // e.g. https://astro.example.com
});
// Everything below is unchanged.
const message = await client.messages.create({
model: "cap:chat.balanced",
max_tokens: 512,
messages: [{ role: "user", content: "Summarise this in one line." }],
});
Already using the OpenAI SDK
from openai import OpenAI
client = OpenAI(
api_key=os.environ["ASTRO_KEY"],
base_url=os.environ["ASTRO_BASE_URL"] + "/v1",
)
completion = client.chat.completions.create(
model="cap:chat.balanced",
messages=[{"role": "user", "content": "Summarise this in one line."}],
)
curl
curl "$ASTRO_BASE_URL/v1/messages" \
-H "x-api-key: $ASTRO_KEY" \
-H "content-type: application/json" \
-d '{
"model": "cap:chat.balanced",
"max_tokens": 128,
"messages": [{ "role": "user", "content": "hello" }]
}'
Rolling back is unsetting one variable. Write your client so the base URL is
process.env.ASTRO_BASE_URL ?? undefined — absent falls through to the vendor
default with no deploy.
Ask for a capability, not a model
cap:chat.balanced rather than claude-haiku-4-5.
A capability is a promise about *what the model must be able to do* — context
window, tool use, vision, residency — resolved to a concrete model at request time.
It survives a model being retired, and it lets an operator move you onto something
cheaper or faster without you shipping anything.
GET /v1/models lists both. Pin a model id when you genuinely need that model; use a
capability the rest of the time.
What comes back
Every response carries headers worth logging:
| Header | Use |
|---|---|
x-astro-request-id | Quote this in a support request. It is the join to the ledger row, the audit entry and the logs |
x-astro-served-model | What actually answered — may differ from what you asked for |
x-astro-cost-micros | Integer micro-USD for this turn |
Errors
Astro returns your vendor's error envelope, so your SDK raises the typed exception
it already raises. The status codes that mean something specific:
| Status | Meaning | Retry? |
|---|---|---|
401 | No key, or a key this deployment does not know | No |
402 | A budget refused this turn. It never reached a provider and cost nothing | No — it will refuse again |
429 | Rate limited | Yes, with backoff |
502 | The provider failed | Yes, if your workload tolerates it |
503 | No credential is wired for a model that could otherwise serve | No — tell an operator |
A 402 carries the numbers to act on:
{
"type": "error",
"error": { "type": "billing_error", "message": "budget exceeded" },
"astro": {
"limit_micros": 25000000,
"committed_micros": 24999100,
"shortfall_micros": 4200
}
}
shortfall_micros is the number that matters. Large means the ceiling is wrong for
your workload; small means a nearly-full budget met an ordinary request and the
window will reset.
Streaming
Set stream: true. You get your vendor's SSE format — Anthropic's named events on
/v1/messages, chat.completion.chunk frames terminated by the literal
data: [DONE] on /v1/chat/completions.
Two behaviours worth knowing:
- A failure before the first byte is an ordinary HTTP error. You can retry it
like any other failed request.
- **A failure after the first byte arrives as an error frame, and the stream
closes.** Astro does not fail over mid-stream: a second model would continue
someone else's sentence.
Surviving an outage
Astro is a dependency. Plan for it being down the way you would any other.
If you already have fallback logic, keep it. The safest adopters put Astro *last*
in an existing fallback chain first, then promote it once it has proven itself.
If you want offline-capable tests, use @astro/sdk, which short-circuits inside
the process with no key and no network:
import { createClient } from "@astro/sdk";
const astro = createClient(); // mode: "auto" — live if credentials resolve, offline otherwise
const res = await astro.chatSoft({
model: "cap:chat.balanced",
messages: [{ role: "user", content: "hello" }],
});
if (!res.ok) {
// Never throws — including on programmer error. Branch on a kind, not a message.
if (res.error.kind === "over_budget") tellTheUser();
else if (res.error.retryable) scheduleRetry();
}
Offline responses are deterministic — the same request gives the same answer on
every machine — so your suite can assert on them. Every response carries source
(live · offline · replay) so you can always tell what answered.
Set ASTRO_STRICT=1 in production. It makes an accidental fall-back to offline
fatal instead of silent. A service quietly answering from a stub because a key was
missing in staging looks healthy while being entirely fictional.
Record real answers once and replay them forever:
import { createClient, fileCassettes } from "@astro/sdk";
const cassettes = fileCassettes("./test/cassettes");
const recording = createClient({ mode: "record", cassettes }); // calls once, saves
const replaying = createClient({ mode: "replay", cassettes }); // never calls
A replay miss is an error, never a live call — otherwise a suite that was meant
to be hermetic quietly starts burning keys.
Rate limits, budgets and cost
Your key is scoped to a project. That project may have a budget with a window
(turn · hour · day · month) and an action when exceeded:
block— refuse with402warn— serve, and reportdowngrade— serve something cheaper
Ask your operator which applies. If you are seeing 402s, the
refusals screen in their console shows exactly how far over each request went.
Checklist before you ship
- [ ] Base URL comes from an env var that can be unset to roll back
- [ ]
x-astro-request-idis logged with your own request id - [ ]
402is handled distinctly from429— one is a ceiling, the other is a wait - [ ]
ASTRO_STRICT=1in production if you use@astro/sdk - [ ] Your CI does not need a network to run its model-touching tests
API reference
The full public surface — every route, every field, with examples — is in the