Rules or a model? The honest answer is a split.

2026-08-26 · essay · rules, models, and the split

Gmail filters and AI triage are usually framed as old-versus-new. They're actually two failure modes in tension: rules fail by going stale, models fail by being unaccountable. Anyone telling you one side wins outright is selling the side they have.

Where filters genuinely win

Where they structurally can't

What a model adds, and what it costs

A language model reads the message and generalizes — new sender, no rule, still a sensible judgment. That solves novelty and staleness in one move. The cost is accountability: ask why a message was buried and you get either nothing or a post-hoc paragraph the model composed after deciding. And the failure is still silent, now with less recourse — there's no rule list to audit.

The split that keeps both halves

The framing error is making one mechanism do both jobs — perceiving the message and deciding what happens. Split them and each side does what it's good at:

How the perception half is chosen isn't a vibe either — we benchmark it across models with a committed eval set anyone can re-run: the six-model bake-off. Rules where rules win, a model where a model wins, and a visible seam between them.

Try Klorn free