Sundr proxy
Sundr sits between your code and your LLM provider. It forwards every call unchanged while it learns a task, then answers the share of calls it can certify within your error budget, at a fraction of the cost. Everything else still goes to your LLM.
The proxy is in design-partner preview. To get a workspace key, email [email protected]. Not ready to route traffic? Run a free audit on exported logs first, or let your coding agent set one up.
Quickstart
Change the base URL, send your workspace key, and name the judgment call with X-Sundr-Task. Your provider key goes where it always does.
OpenAI Python SDK
client = OpenAI(
base_url="https://proxy.sundr.ai/openai/v1",
default_headers={"X-Sundr-Key": os.environ["SUNDR_KEY"]},
)
result = client.chat.completions.create(
model="gpt-4.1-mini",
messages=[{"role": "system", "content": ROUTER_PROMPT},
{"role": "user", "content": ticket_text}],
response_format={"type": "json_schema", "json_schema": {"name": "route", "schema": {
"type": "object", "required": ["queue"],
"properties": {"queue": {"enum": ["billing", "bug", "how-to", "other"]}}}}},
extra_headers={"X-Sundr-Task": "ticket-router"},
)
Anthropic Python SDK
client = Anthropic(
base_url="https://proxy.sundr.ai/anthropic",
default_headers={"X-Sundr-Key": os.environ["SUNDR_KEY"]},
)
client.messages.create(..., extra_headers={"X-Sundr-Task": "ticket-router"})
TypeScript
const openai = new OpenAI({ baseURL: "https://proxy.sundr.ai/openai/v1", defaultHeaders: { "X-Sundr-Key": process.env.SUNDR_KEY }, }); await openai.chat.completions.create(body, { headers: { "X-Sundr-Task": "ticket-router" } });
Supported endpoints: POST /openai/v1/chat/completions and POST /anthropic/v1/messages.
SDKs
Optional helpers do the setup above for you. They read SUNDR_KEY and return the SDK's own client or options, so the rest of your code stays the same. The packages are published to PyPI and npm with the public launch; design partners get them from us directly until then.
# Python: sundr.client from sundr.client import openai_client, task client = openai_client() # or anthropic_client(), async_openai_client(), async_anthropic_client() client.chat.completions.create(..., extra_headers=task("ticket-router"))
// TypeScript: @sundr/sdk import { openaiOptions, task } from "@sundr/sdk"; const openai = new OpenAI(openaiOptions()); // or new Anthropic(anthropicOptions()) await openai.chat.completions.create(body, task("ticket-router"));
Both packages also have served_by / servedBy, which reads x-sundr-served-by from a raw response, and a Control client for the control API.
What Sundr answers
A call is a judgment call when it carries X-Sundr-Task and asks for structured output with a fixed set of answers:
- a JSON schema in
response_format(OpenAI) oroutput_config.format(Anthropic), or a forced tool call with a parameter schema, and - a field in that schema with a fixed set of values: an
enum, a boolean, or a bounded integer. That field is the verdict. Up to 255 labels.
Sundr finds the verdict field automatically. If several fields qualify, it prefers one named like a verdict (label, category, decision, is_…); if it still can't tell, the task stays in shadow until we set the field with you.
Use one task name per decision, not per prompt version. Everything else passes straight through untouched: calls without X-Sundr-Task, streaming calls, and calls whose output isn't a fixed set of labels.
Responses
When Sundr answers, the response has the same shape your code already parses: the same API format, a body that satisfies your schema, and the same channel (message text or the forced tool call). Every response carries x-sundr-served-by: sundr or llm.
Fields other than the verdict get schema-valid defaults: confidence-like numbers get Sundr's score, and rationale-like strings say the call was answered by Sundr. If your code depends on a free-text explanation, keep that call on your LLM. Usage reports zero tokens, because none were billed.
Task lifecycle
- Shadow
- Every call goes to your LLM. Sundr logs the input and the verdict it returned.
- Certified
- From at least 2,000 shadow calls, Sundr has trained a model and certified a threshold for your error budget (1% unless agreed otherwise). Still nothing is answered by Sundr.
- Takeover
- You start it. Calls Sundr scores above the threshold are answered by Sundr; the rest go to your LLM.
- Pass-through
- Everything goes to your LLM again. Entered when you pause takeover, or automatically when the live monitor finds the budget broken.
The certificate
Sundr splits the shadow log into training and held-out calibration calls, then walks thresholds from strict to loose with a binomial test at each (Learn then Test), stopping at the first it can't vouch for. The certificate states the threshold, the share of calls it covers, and an upper bound on disagreement with your LLM that holds with 95% confidence, provided new traffic resembles the logged sample. Reliable certification takes about 120 ÷ budget calibration calls: roughly 12,000 for 1%, 6,000 for 2%.
The guarantee is about agreement with your LLM, not about ground truth. If your LLM is wrong, Sundr agreeing with it is not a fix.
Live monitoring
During takeover, a random 2–10% of the calls Sundr answers are also sent to your LLM, more often near the threshold. Every ten minutes Sundr re-estimates the error on the last seven days of answered calls (active statistical inference, an unbiased estimate with a 95% upper bound). If the audits show the budget is broken, the task moves to pass-through on its own and the dashboard says why. Audit calls are billed by your provider as normal and are excluded from savings.
Coverage mode
Coverage mode is a second opinion on every call your LLM answers. The LLM still answers; Sundr scores the same call in the background and logs its own verdict beside the LLM's. The dashboard shows how often they agree and lists the calls where Sundr disagreed at or above its certified threshold. Those are worth reading: they are often LLM mistakes, prompt regressions or a shift in traffic. It needs a certified task, works in any state, and never changes a response or adds latency. Turn it on from the dashboard or the control API.
Control API
Send X-Sundr-Key to https://proxy.sundr.ai.
# tasks, states, certificates and live error curl https://proxy.sundr.ai/sundr/v1/tasks -H "X-Sundr-Key: $SUNDR_KEY" # coverage mode on or off curl -X POST https://proxy.sundr.ai/sundr/v1/tasks/ticket-router/coverage \ -H "X-Sundr-Key: $SUNDR_KEY" -H "content-type: application/json" -d '{"enabled": true}' # pause takeover: everything goes to your LLM curl -X POST https://proxy.sundr.ai/sundr/v1/tasks/ticket-router/state \ -H "X-Sundr-Key: $SUNDR_KEY" -H "content-type: application/json" \ -d '{"state": "passthrough"}'
States: shadow, takeover (needs a certificate) and passthrough. The same controls are on the dashboard.
Data handling
- Your provider key passes through on each request. It is never stored or logged.
- For judgment calls Sundr stores the input text, the verdict, token counts, model and latency, to train and audit the task. Logged calls are deleted after 30 days; certificates and task models are kept.
- To answer a call, its text is sent to decision-model providers (TypeSafe via OpenRouter, Cloudflare Workers AI). Question drafting sends a sample to Anthropic. Nothing is used to train models, and nothing is shared across customers except question templates.
- Hosted on Fly.io in the United States. See the privacy policy.
Pricing
Shadow mode and the audit are free. In takeover, Sundr charges 25% of measured savings: for each call it answers without an audit, what your LLM would have charged for it, computed from your own logged token counts and list prices. If it saves nothing, it costs nothing. Coverage mode is $250 per million calls scored. The dashboard shows savings and fees as they accrue, and the savings ledger downloads as CSV, one row per task per day.