sundr Savings Audit

Phishing triage · sample report

This is a real Savings Audit run on the public jev-phishing-bench dataset (1,934 emails), with the dataset's labels standing in for an LLM's past verdicts, and an assumed 2,000,000 calls a month at $0.004 each. It tested decision models only (Clef-flash and Jev). New audits also test smaller LLMs as candidates, including a plain switch to Claude Haiku 5.5, and show certified coverage and savings for each; otherwise your report looks like this, computed on your own logged calls. Fees below are at the Takeover plan's 20% of savings. Run yours free →

1,934 logged calls · clef-flash (clef-flash) · jev (typesafe/jev-1.13) · 2026-10-09T07:33:56+00:00
At a 1% error budget, Sundr could answer 99% of these calls, saving you $6,176 a month after fees.
Certified with 95% confidence on 774 held-out calls. Every other call would still go to your current model.
99.5%
agreement with your judge, held-out calls
79.3%
asking the decision model one direct question
$0.3011
decision-model cost per 1,000 calls
515 ms
median latency (p95 810 ms)

Certified coverage by error budget

0%25%50%75%100%Sundr: 94% at a 0.5% budget94%One question: 0% at a 0.5% budget0%0.5% budgetSundr: 99% at a 1.0% budget99%One question: 0% at a 1.0% budget0%1.0% budgetSundr: 100% at a 2.0% budget100%One question: 0% at a 2.0% budget0%2.0% budgetSundr: 100% at a 5.0% budget100%One question: 0% at a 5.0% budget0%5.0% budget
Coverage is the share of calls Sundr answers itself. “One question” asks the decision model for the verdict directly, without splitting it.

By budget

Error budgetCalls answeredCertified boundOne questionSavings / monthFeeNet to you
0.5% 94% ≤ 0.41% 0% $7,279$1,456$5,823
1% 99% ≤ 0.61% 0% $7,720$1,544$6,176
2% 100% ≤ 1.18% 0% $7,760$1,552$6,208
5% 100% ≤ 1.18% 0% $7,760$1,552$6,208
Assumes 2,000,000 calls a month at $0.0040 each, the Takeover plan's 20% fee on savings, and about 3% of answered calls still sent to your model as an ongoing audit.

Data needed

Certifying a budget reliably takes roughly 120 ÷ budget held-out calls: about 12,000 for 1% and 6,000 for 2%. This audit held out 774. More logged calls, or a few days of shadow mode, raise certified coverage at tight budgets.

Questions Sundr asks

17 kept of 17 drafted · each answered by every decision model, choice questions in two option orders

NameTypeQuestion
verdictchoiceYou are an email security assistant. Read the email (sender, from address, subject, body, link text and link URL) and decide whether it is a phishing attempt whose link the user should not click, or a legitimate email.
phishing · legitimate
verdict_altchoiceWhich label best describes this text? Task: You are an email security assistant. Read the email (sender, from address, subject, body, link text and link URL) and decide whether it is a phishing attempt whose link the user should not click, or a legitimate email.
phishing · legitimate
verdict_mirrornoulThe correct label for this text is 'phishing'. Task: You are an email security assistant. Read the email (sender, from address, subject, body, link text and link URL) and decide whether it is a phishing attempt whose link the user should not click, or a legitimate email.
sig_sender_domain_mismatchnoulThe sender's display name or claimed organisation belongs to a different organisation than the domain of the 'from' address.
sig_link_domain_mismatchnoulThe link's destination domain belongs to a different organisation than the sender's email domain.
sig_free_hostingnoulThe link points to a free web-hosting, file-sharing, form-builder or URL-shortening service.
sig_odd_urlnoulThe link URL uses a raw IP address, a long random-looking path, or a misspelled brand name.
sig_display_text_mismatchnoulThe link's display text names a different website than the URL it actually points to.
sig_lurenoulThe email offers a prize, refund, payment or reward.
sig_urgencynoulThe email pressures the reader to act immediately or warns that an account will be suspended.
sig_credentialsnoulThe email asks the reader to sign in, verify an account, or confirm personal or payment details.
sig_free_mail_businessnoulThe sender claims to represent a business while writing from a free consumer email provider or an automatically generated address.
sig_generic_greetingnoulThe email greets the reader generically, such as 'Dear customer', instead of by name.
sig_brand_impersonationnoulThe email claims to come from a well-known brand, bank, delivery company or online service.
sig_routine_businessnoulThe email reads like routine correspondence between colleagues or within an established business relationship.
sig_reputable_linknoulThe link points to a widely known, reputable domain such as a major company's official website.
suspicionscoreHow suspicious would a careful security analyst find this email overall?
routine · slightly unusual · moderately suspicious · highly suspicious

How this was measured

1,160 calls trained a small logistic-regression combiner over the decision models' answers. Its threshold was certified on 774 separate calls with Learn then Test (fixed-sequence binomial tests, 95% confidence). The bound holds for future traffic that resembles these logs; in production a live audit slice re-checks it and hands traffic back to your model if it slips.

Labels seen: legitimate (1,000) · phishing (934)

Next: a 30-day free shadow run. One base-URL change; nothing changes for your users until a budget is certified on live traffic.[email protected] →