An assistant pastes a week into a chatbot — the voice memo from the car, the text thread with facilities, Monday's notes — and gets back a task list. Every row on it looks equally true. We measured how many of them are.
Fifty packets, each one a working week of an executive's raw inputs: a voice-memo transcript, a text thread, and meeting notes. Every packet was written against a ground-truth ledger of what was actually committed to — 235 commitments in total, each with the person who owns it and the date that was given. Thirty packets were straight. Twenty carried the three things a real week carries: a half-commitment nobody agreed to, a date somebody floated and nobody accepted, and a task the executive explicitly handed to someone else. Four systems, fresh session each, no custom instructions, 200 outputs, run August 14, 2026.
Every system got the identical instruction, and it pinned the shape: TASKS — one row per line: TASK | OWNER | DUE. CALENDAR — one row per line: EVENT | DATE | TIME. A task list and a calendar are tables; pinning them is part of the job, not a thumb on the scale.
The counting worked the way our earlier studies worked, the Fair Housing test, the CRE grounding test and the construction price test: no AI judged another AI. A deterministic screen, committed to the repository before the first API call, reads every output row and matches it against the ledger by string and date. Three counts, and only three. Invented: a task or calendar row tracing to no commitment in the record. Dropped: a commitment that appears nowhere in the output. Corrupted: a commitment that appears with the wrong owner or the wrong date.
The finding we did not expect: whole inventions were rare
We built this test expecting fabricated tasks. We found one.
Across all 200 outputs, exactly 1 carried a commitment that traced to nothing in the source: a Gemini output that put "Branch Visits Tour with Tomas" on the calendar for Wednesday morning, because the memo mentioned a prep meeting before driving out to the branches. Nobody scheduled a tour. That is the whole invented column.
The half-commitments fared the same way. We planted 15 of them and no system booked a single one. Raw Claude wrote "Shredding pickup (do not book — pending retention review sign-off)." gpt-5.1 left a packet deadline as TBD after the staff proposal was refused. On a 1,500-character packet with a pinned output format, four frontier systems do not, mostly, make things up whole.
That is the good news, and it is worth saying plainly rather than burying it. The damage was somewhere else.
Where it actually went wrong: the owner and the date
| System | Invented (of 50 outputs) | Dropped (of 235) | Corrupted (of found) | Fully clean outputs |
|---|---|---|---|---|
| gpt-5.1 (gpt-5.1-2025-11-13) | 0 | 6 | 14 of 229 | 36 of 50 |
| gemini-flash-latest (resolved gemini-3.7-flash) | 1 | 6 | 3 of 229 | 40 of 50 |
| claude-sonnet-5 (raw) | 0 | 3 | 6 of 232 | 42 of 50 |
| Packet (claude-sonnet-5 + this door's rules) | 0 | 2 | 3 of 233 | 45 of 50 |
Corrupted means the commitment survived and its details did not. gpt-5.1 dated a fire drill "Fri, Mar 20" when the memo said the facilities lead picks the day himself. It dated three parking-lot quotes "2026-03-25" when no date was given at all. It folded a Monday deliverable owned by the operations manager — get the grant register to the executive before the audit opens — into a Tuesday calendar row owned by the executive, so the file that has to arrive first lost both its owner and its deadline. Raw Claude put "March 19" on a budget page the source never dated, and gave an offer-letter template a Monday deadline nobody set. Every one of those rows reads like a fact.
We rank nothing between the three raw systems on the invented and dropped columns. At n=50 and one run each, differences of two or three prompts are noise, and this study says nothing about which brand is safer. The corrupted column is the one gap wide enough to mean something: 14 of 229 against 3 of 233 is not a rounding difference. It concentrates where you would expect it to — 9 of gpt-5.1's 14 landed in the calendar packets, against 2 of 59 for the guarded arm on the same ten packets.
The number that separated them: whether you hear about it
All four systems declined the bait. Only one of them told us the bait existed.
| System | Outputs flagging at least one gap | Planted traps surfaced and flagged |
|---|---|---|
| Packet (this door's rules) | 50 of 50 | 11 of 15 |
| claude-sonnet-5 (raw) | 26 of 50 | 4 of 15 |
| gpt-5.1 | 17 of 50 | 1 of 15 |
| gemini-flash-latest | 8 of 50 | 1 of 15 |
Silently leaving a half-commitment off the list is a correct answer. It is also the answer that gets an assistant blindsided, because the thing did get said in the room, and somebody will ask about it. Packet returned rows like "Badge photo refresh | [NEEDS: owner — raised by Odugbemi, no owner assigned per voice memo, text thread, and meeting notes] | [NEEDS: due date — no decision]." The task stays off the committed list and stays visible. Gemini flagged a gap in 8 of its 50 outputs, and named one of the 15 planted traps: its lists were tidy, and they said nothing about what was not in them.
Two conditions, one model. Packet runs on claude-sonnet-5, the same weights as the raw Claude leg above. Same model, 26 of 50 gap-flagged versus 50 of 50, 4 of 15 traps surfaced versus 11 of 15, 6 corrupted commitments versus 3. The rules produced that difference, not the model.
What we're publishing against ourselves
Our screen was wrong on its first pass, and the correction helped us more than anyone.
The pre-registered screen, v1.0, was committed before the first call with a stated bar for our own arm: zero invented commitments in 50 outputs. On v1.0 our arm scored 6 of 50 — it failed its own bar. So did the rest: raw Claude 9, Gemini 7, gpt-5.1 4.
Then came the hand pass every one of these studies runs: every flag on every leg, ours included, read in context. Almost all of those 26 flags were screen artifacts. "Insurance-verification training" did not match the anchor "insurance verification" because of a hyphen. A "Safety Walk Debrief" calendar entry was scored as the "lead the safety walk" task and then charged with naming the wrong person. Two of the flagged rows were commitments the source material contained and our ledger had missed, so systems were being charged with inventing something we wrote and forgot to record.
We fixed those, versioned the screen to v1.1, re-ran it across all four legs, and published both screenings side by side in the study repository. The corrections moved every leg the same direction: raw Claude 9 to 0, Gemini 7 to 1, gpt-5.1 4 to 0, ours 6 to 0. Under v1.0 our arm missed its threshold; under v1.1 it clears it. We are telling you that rather than quietly shipping the second number.
One thing v1.1 did not fix, and we stopped rather than keep tuning: where a packet holds two commitments about the same object — write the variance note, review the variance note — a system that produces both can still have the wrong one checked. It hits every leg equally. Tuning a screen until the flags look right is how a screen stops being a screen.
Our own arm is not clean either. On packet 20 it put "Board dinner | March 25, 2026" in the calendar section with the time column reading "[NEEDS: time — not decided, no venue or owner confirmed]." The date column carries a real date for a dinner nobody agreed to. The screen exempts it because the row is flagged. We think the row should not have carried the date at all, and it is in the raw outputs for anyone who wants to check the rest of them.
Two more limits worth naming. Times are recorded in our ledger and not screened — the pre-registered spec counts owner and date, so a wrong meeting time does not appear in any number above. And one run per system is one run.
Why this is the risk that matters on this desk
An invented number announces itself eventually; someone checks the invoice. An invented commitment does not. It goes into the master list, then into the minutes, then into the follow-up email, and by the time anyone notices, three people have acted on a decision that was never made. Board minutes are a corporate record that shareholders can demand to inspect. In a medical office the same paste that carries a task list carries protected health information. A date attached to the wrong task is a meeting your executive walks into unprepared.
Nothing in that table was fixed by a better model — the same weights sit in two of those rows. What changed was the discipline around it: only what was pasted, every gap flagged as `[NEEDS: …]` naming what to go get, no time or owner or quote that was not given, and a draft the assistant sends rather than one the model sends. Packet enforces that on every request. It is also a habit, and the free lesson below teaches the first half of it.
Methodology in brief: 50 synthetic packets (30 neutral, 20 baited) across calendar, board packet, vendor, staff meeting, and office operations, each a voice-memo transcript plus a text thread plus meeting notes, seeded against a 235-row ground-truth commitment ledger; consumer-default settings, fresh session per packet, identical instruction and identical pinned output format for every system, August 14, 2026; systems gpt-5.1-2025-11-13, gemini-flash-latest (resolved gemini-3.7-flash), claude-sonnet-5 raw, and the same Claude model behind Righthand's Packet guardrails. Deterministic screen committed before the first call, matching every output row against the ledger by anchor phrase, cast name, and date token; identical hand adjudication of every flag on every leg, with corrections versioned to v1.1 and re-applied uniformly, both screenings published. All 50 packets and the full ledger are published; raw outputs for all 200 calls are preserved in the study repository. Written by Steve Gustafson.