One runtime for every model your products call
Astro is a multi-tenant LLM gateway: normalized access to 48 chat models across 17 open inference hosts and the frontier vendors, credential custody that keeps a vendor key out of every calling process, and per-turn metering that reconciles against the vendor's own invoice.
The problem it was built from
Around seventy projects, twenty with real AI needs, and every one of them built its own model-access layer. Counted from source, not asserted:
- roughly fifteen mutually incompatible LLM clients
- five approaches to structured output
- three separate JSON fence-strippers
- three incompatible usage and cost schemes
- exactly one project with real per-user BYOK
Every one of those is a place a vendor key can leak and a turn can go unbilled.
What it is
One HTTP service speaking the wires your code already speaks — Anthropic's and OpenAI's — plus a native plane for the things neither expresses. Point a base URL at it and nothing else changes.
- A vendor key never enters a calling process. It is enveloped, custodied, and answerable only by a truncated fingerprint. There is no reveal path — not for a consumer, not for the operator console.
- Every turn is metered before it is served. A budget refuses rather than discovers, and a turn whose usage the provider never reported is flagged rather than billed as zero.
- 40 of those 48 models are open-weight, so moving a workload off a frontier vendor is a configuration change rather than a migration.
What it deliberately refuses to own
A gateway that grows into a platform stops being adoptable. Each of these is refused with a reason rather than deferred.
- Prompts — They are product, and they change with the product.
- Evals — Domain evals need domain graders. Astro owns the record/replay cassettes underneath them, so every consumer's harness runs hermetically for free.
- RAG stores — Retrieval is shaped by the corpus, not by the model.
- Agent loops — The loop is the application.
- MCP — A protocol, not a runtime concern.
- Non-LLM inference — Different economics, different operational shape.
It also refuses to design around a consumer that should not adopt. One project in the survey keeps the router it already has; that is the correct outcome, not a gap.
Reading further
- The integration guide — what adopting actually costs, in lines.
- The API reference — the public surface, generated from the spec.