Klorn Download

Which LLM should classify your email? I ran ten, three times each.

2026-08-26 · re-measured 2026-09-04 after a reader’s objection · provider-default temperature · every number re-runs with pnpm eval:judge

Every AI email tool picks a model for you and never shows its work. Klorn’s eval set is committed to the repo and gated in CI, so instead of an opinion, here is a measurement: ten current models, one classification job, identical prompt, identical decision rule, three runs each.

The job is narrow on purpose. The model does not pick the folder or the priority. It scores four features per email — confidence, sender trust, reversibility, urgency — and a deterministic ~200-line rule maps those numbers to exactly one of five lanes. That split is what makes this benchmark possible at all: the model’s task is small enough to measure.

Why this post changed. The first version (2026-08-26) ran six models once each and ranked them. A reader pointed out that on 56 items every 1.8 points is a single email, and that his own two-runs-an-hour-apart test had swapped first and last place with nothing wrong. He was right to object, so I re-ran everything three times — and found two more problems of my own: the six models were the product’s chat catalog, pinned in July, which had silently missed four newer models; and the original footnote said “temperature-0”, which was false — the judge sets no temperature and runs at the provider default. Both are corrected below.

ModelRunsCorrect / 56RangeUrgent / 13Gate$/M in
gpt-5.4356 · 56 · 56100.013 · 13 · 133/3$2.50
gemini-3.5-flash355 · 55 · 5598.213 · 13 · 133/3$1.50
gpt-5.6-terra255 · 5598.213 · 132/2$2.00
grok-4.6 *15598.2121/1$2.00
gemini-3.7-flash355 · 55 · 5496.4–98.213 · 13 · 123/3$0.75
gpt-5.6-luna253 · 5594.6–98.212 · 122/2$0.20
gemini-2.5-flash (default pin)354 · 54 · 5496.413 · 13 · 133/3$0.30
grok-4.3254 · 5496.412 · 122/2$1.25
claude-opus-4.8250 · 4987.5–89.39 · 80/2$5.00
claude-sonnet-5248 · 4580.4–85.77 · 50/2$2.00

Committed 56-email gate set, --context=fixture, JUDGE_INCLUDE_BODY=true, no temperature set (provider default) — the same configuration production runs. Gate floors: overall ≥80%, urgent recall ≥90%, silenced precision ≥90%. 23 of 30 planned runs are shown: runs in which any item fell back to the keyword path (a provider error, not a model verdict) were discarded rather than averaged. * grok-4.6 lost two runs to upstream timeouts of over 90 minutes each; its single clean run is reported as such. The sweep also shared an API key with the hosted service and briefly rate-limited it — that is written up in the ops repo, and the keys are separate now.

Five things in this table matter more than the ranking — and the first one is that there mostly isn’t a ranking.

1. Tiers, not ranks

The reader was right about the middle: four models sit at 98.2 and nothing in this data orders them — a one-email gap on 56 items is not a gap. He was wrong about the top and the bottom. gpt-5.4 scored 56/56 on all three runs; both Anthropic models failed the urgent-recall gate on every run they had. Three runs turn “first place” into “stable top tier” and “fifth place” into “stable failing tier”, and that is the shape the ranking should have had from the start.

2. Price does not order the table — more so, not less

The $5.00/M model never passes the gate. The $0.20/M model reaches 98.2 in one run and 94.6 in the other, landing in the same tier as models costing seven to ten times more. The spread between the top tier and the failing tier is 15–20 points depending on the run, which is far above the one-to-two-email noise floor the middle of the table lives in. If you chose your model because it tops a general benchmark, you chose on a number that has nothing to do with this job.

3. The interesting failure has a mechanism — and it reproduced

In the first run claude-sonnet-5 missed eight urgent emails, and seven of the eight failed on confidence, not urgency: it scored urgency 0.80–1.00 — correct — and reported confidence 0.55–0.60, under the 0.70 bar the rule requires before interrupting you. It read the mail right and declined to say it was sure. Across the re-runs it caught 5 to 7 of 13 urgent items; the mechanism did not change, only the count.

I can tell you that because the threshold is a number in a file. In a “model picks the label” design the same result reads as “this model is worse at email”, which is both wrong and unfixable.

4. Stability is a property, and it is worth paying for

At provider-default temperature, seven of the nine models with repeat runs moved by at most one email between runs. The two exceptions split: gpt-5.6-luna moved by two and still passed the gate both times; claude-sonnet-5 moved by three and failed it both times. grok-4.6 has a single clean run — two were lost to upstream timeouts — so it has no spread to report. This is why gemini-2.5-flash stays the default pin: not because 96.4 is the best number in the table, but because it was 54/56 three times at $0.30. For a judge that runs on every inbound email, a model that is a little worse and never surprises you beats one that is a little better and sometimes does.

5. The floor held for every model, every run

Across 23 runs, ten models and a 20-point accuracy spread, the number of urgent emails classified as hide-forever was zero — every miss degraded to the visible queue. That invariant is a unit test, not a hope, and it is the only line in this post that will still be true when the next model generation ships.

Which is the point. In June this set said a cheap model beat the frontier. In August one run said the opposite and I retired the claim. A reader then noted, correctly, that a retraction on n=1 has exactly the standing of the claim it replaces. So: n=3, the middle of the table is now honestly a tie, and the two claims that survived are the ones that were about the whole table rather than any adjacent pair. The benchmark’s job is not to crown a model — it is to make the choice re-runnable in one command when the ground shifts again:

JUDGE_MODEL=<any-openrouter-id> pnpm eval:judge

The eval set, the rule, the floors and the receipts pipeline are AGPL-3.0: github.com/k08200/klorn. Or skip the setup and point Klorn at your own inbox:

Try Klorn free