Two weeks ago I published the build log of an email-triage agent running on the cheapest desktop Apple sells. It ended on a promise: the next improvement would take the same machine further, with nothing bought. This is the follow-up — what got faster, what got more precise, and what the project is turning into.
Performance: three times the mail, same machine
The full 206-email test set now runs in under 17 minutes, against 51 minutes for the build in the first article. None of it came from a bigger model or a new machine. Three unglamorous changes did the work:
- The model stays loaded. The model server unloaded it after twenty idle minutes, and the next email paid about five seconds to bring it back. On a live inbox, where mail arrives in bursts, that was most emails.
- Two emails at a time. The Mac mini turned out to serve two requests side by side with little loss on each, so the queue now drains through two worker processes. The ceiling roughly doubles for free.
- Less text per email. Tighter instructions and a shorter answer format cut what the model reads and writes per message by about a sixth — and on this hardware, the wait follows the text.
One idea was measured and thrown out on the way: abbreviating the answer's field names to save output tokens. It was faster, and it failed the release gate — two real inquiries binned. For a model this size the name of a field is part of its definition. The gate did its job.
Accuracy: more precise sorting, still nothing lost
The first version answered one question per email — reply, read, archive or bin. The current one files each message into one of thirteen folders: Action, Receipts, Newsletters, Cold outreach, Spam, Notices for the server and website alerts a business runs on, and a set of Reference folders for clients, finance, admin, accounts and personal mail. A harder job, with the same rules: never bin a legitimate email, never miss one that needs a reply.
The new number that matters is the third one. The synthetic test set is still the gate, but there is now a second one: forty-five real emails from my own inbox, labelled by hand and never published. The current build files 41 of them where I would have, and the four misses are judgement calls — a cold pitch filed as something to act on, a forwarded personal note filed as work. Nothing real went to the bin.
Most of the gain came the same way as before: from corrections. Every time I moved an email the agent had filed, that correction was recorded, and patterns in them became rule changes — promotions from companies I have an account with are newsletters, not reference; a stranger pitching services from a free email address is cold outreach, even phrased as a question. The next layer, being tested now, learns those patterns per sender on its own instead of waiting for me to write them down.
| First article (Sep 17) | Today (Oct 1) | |
|---|---|---|
| Time per email | 19 s, one at a time | 9.5 s each, two at a time |
| Ceiling, one Mac mini | ~4,500 a day | ~17,000 a day |
| What it decides | 4 verdicts | 13 folders, plus a verdict |
| Needs a reply, P / R | 0.97 / 0.96 | 0.93 / 0.95 |
| Legitimate emails binned | 0 | 0 |
| Real inbox check | — | 41 of 45 filed right |
The reply-needed scores moved by a point or two in both directions as the job grew from four answers to thirteen. They stay well above the gate, which is the point of having one.
From an agent to a mail app
A triage agent that only sorts still leaves you in someone else's mail client, looking at its folders. So the project is becoming the client itself: one app over every mailbox you have, with the AI inside it rather than bolted on.
- One inbox across accounts. Every folder of every mailbox is synced, and each message stays owned by its own account — a reply goes out from the address it was sent to.
- Sending and drafts. Send with a ten-second undo, drafts kept in step with the mail server so they show up on any other device.
- AI reply drafts, never sent on their own. On request, the local model writes a reply from the whole thread and examples of how you write. You edit and send it; it never leaves by itself.
- Search in plain English. "Invoices from the printer last spring" works without learning a query syntax.
- Triage as an option. The app works as a normal mail client first. Sorting, auto-filing and learning are switched on per company, and each one separately.
The mail engine is done; the app is being built now. Syncing, sending, drafts and the AI features run today on the server side. The app on top — one codebase for iPhone, Android and the web — is the current piece of work.
What this means for an SME
The pattern from the first article held. A tripling of capacity came from engineering — keep the model warm, run two at a time, send less text — not from spending. Accuracy improved because real corrections were recorded and turned into tested rules, and a release gate threw out the changes that looked good but lost mail.
That is the part worth paying for. The machine is cheap, the model is free; making them dependable enough to trust with your inbox is the work.