Your 1% threshold is really 1.5%
October 9, 2026
Decision models like Jev and Clef return calibrated probabilities, and the obvious way to use them is a confidence threshold. Answer automatically above it, and send everything below it to a person or a bigger model. Pick the threshold on a labelled sample so that the error on that sample is under your budget, say 1%, and ship.
That threshold will break your budget about half the time.
The overshoot is not bad luck
Vals AI’s independent evaluation of Jev did exactly this. They tuned a threshold on 30% of their items for a 1% error budget, then measured 1.6% error on the other 70%. Six of the twelve systems they tested overshot the budget they had been tuned for.
It isn’t a quirk of those models. It’s what happens when you choose the threshold on the same sample you measure it on. The procedure keeps loosening the threshold for as long as the observed error stays under 1%, so it stops exactly where luck made the sample look best. On new traffic, the luck runs out.
We simulated it. The model is deliberately a little overconfident: true error is 1.3× what its scores imply, which is mild by the standards of the calibration critiques published since Jev launched. For each calibration size we drew 300 independent samples and tuned a threshold on each.
| Calibration items | Naive threshold: over budget | Naive: mean true error | Certified: over budget | Certified: share of traffic answered |
|---|---|---|---|---|
| 120 | 56% | 1.49% | 0% | 0% |
| 400 | 62% | 1.21% | 0% | 0% |
| 1,000 | 55% | 1.08% | 1% | 4% |
| 3,000 | 58% | 1.06% | 2% | 13% |
| 12,000 | 50% | 1.00% | 3% | 38% |
The best threshold possible on this population answers 49% of traffic at exactly 1% error.
Two things stand out:
- More data doesn’t fix the naive method. It just centres the error on the budget, so it overshoots half the time instead of most of the time. “Under 1% on our test set” is a coin flip on production.
- 120 items, the size Vals used, averages 1.49% true error. That’s close to their measured 1.6%.
What a guarantee costs
The certified column uses Learn then Test (Angelopoulos, Bates, Candès, Jordan and Lei). Walk the thresholds from strict to loose. At each one, run a binomial test of “the error here is above 1%”, and stop at the first threshold where you can’t reject it. The result holds on new traffic with 95% confidence, whatever the model’s calibration, as long as new traffic resembles the sample.
The price is coverage, and it’s paid in data. With zero errors you need at least 300 accepted items to certify 1% at 95% confidence, because the rule of three gives 3/300. In practice you need around 120 ÷ budget calibration items, about 12,000 for a 1% budget, before the certified threshold gets close to the ideal one. Below about 1,000 items, the honest answer is “you can’t promise 1% yet”.
That’s the uncomfortable truth behind every confidence threshold: the number you can promise depends on how much evidence you have, not just on how good the model is.
What this means if you’re replacing an LLM judge
- Don’t tune and measure on the same sample. If you must hand-tune, at least hold out a second sample and report its error with a confidence interval.
- Size the calibration set for the budget you want. A few hundred labels supports 3–5%, not 1%.
- Log enough traffic before you automate. Logs are free labels when you’re replacing an existing LLM judge: its past verdicts are what you need to agree with.
- Keep checking after launch. A threshold certified on last month’s traffic says nothing about a new product line. Send a small, uncertainty-weighted slice of automated calls to the reference model, and estimate live error with a valid interval.
Sundr does all four. It’s free to try on your own logs: upload a week of judgment calls and it tells you how many it could have answered, at a certified error rate. Run a Savings Audit →
Method: 300 simulated calibration samples per row, seed 2026. The simulation runs on a laptop in under a minute.