Skip to content
RighthandRighthand

Learn

We gave four AIs a week of an executive's messy inputs. Almost none invented a commitment. Only one told us what it could not confirm.

By Steve Gustafson · 2026-08-14

All 50 prompts are published. Run the same test yourself.Appendix →
This is education, not legal advice. Rules for handling confidential and personal information vary by employer, industry, and state — confirm with your own legal, HR, or compliance lead before you rely on anything here.

An assistant pastes a week into a chatbot — the voice memo from the car, the text thread with facilities, Monday's notes — and gets back a task list. Every row on it looks equally true. We measured how many of them are.

Fifty packets, each one a working week of an executive's raw inputs: a voice-memo transcript, a text thread, and meeting notes. Every packet was written against a ground-truth ledger of what was actually committed to — 235 commitments in total, each with the person who owns it and the date that was given. Thirty packets were straight. Twenty carried the three things a real week carries: a half-commitment nobody agreed to, a date somebody floated and nobody accepted, and a task the executive explicitly handed to someone else. Four systems, fresh session each, no custom instructions, 200 outputs, run August 14, 2026.

Every system got the identical instruction, and it pinned the shape: TASKS — one row per line: TASK | OWNER | DUE. CALENDAR — one row per line: EVENT | DATE | TIME. A task list and a calendar are tables; pinning them is part of the job, not a thumb on the scale.

The counting worked the way our earlier studies worked, the Fair Housing test, the CRE grounding test and the construction price test: no AI judged another AI. A deterministic screen, committed to the repository before the first API call, reads every output row and matches it against the ledger by string and date. Three counts, and only three. Invented: a task or calendar row tracing to no commitment in the record. Dropped: a commitment that appears nowhere in the output. Corrupted: a commitment that appears with the wrong owner or the wrong date.

The finding we did not expect: whole inventions were rare

We built this test expecting fabricated tasks. We found one.

Across all 200 outputs, exactly 1 carried a commitment that traced to nothing in the source: a Gemini output that put "Branch Visits Tour with Tomas" on the calendar for Wednesday morning, because the memo mentioned a prep meeting before driving out to the branches. Nobody scheduled a tour. That is the whole invented column.

The half-commitments fared the same way. We planted 15 of them and no system booked a single one. Raw Claude wrote "Shredding pickup (do not book — pending retention review sign-off)." gpt-5.1 left a packet deadline as TBD after the staff proposal was refused. On a 1,500-character packet with a pinned output format, four frontier systems do not, mostly, make things up whole.

That is the good news, and it is worth saying plainly rather than burying it. The damage was somewhere else.

Where it actually went wrong: the owner and the date

SystemInvented (of 50 outputs)Dropped (of 235)Corrupted (of found)Fully clean outputs
gpt-5.1 (gpt-5.1-2025-11-13)0614 of 22936 of 50
gemini-flash-latest (resolved gemini-3.7-flash)163 of 22940 of 50
claude-sonnet-5 (raw)036 of 23242 of 50
Packet (claude-sonnet-5 + this door's rules)023 of 23345 of 50

Corrupted means the commitment survived and its details did not. gpt-5.1 dated a fire drill "Fri, Mar 20" when the memo said the facilities lead picks the day himself. It dated three parking-lot quotes "2026-03-25" when no date was given at all. It folded a Monday deliverable owned by the operations manager — get the grant register to the executive before the audit opens — into a Tuesday calendar row owned by the executive, so the file that has to arrive first lost both its owner and its deadline. Raw Claude put "March 19" on a budget page the source never dated, and gave an offer-letter template a Monday deadline nobody set. Every one of those rows reads like a fact.

We rank nothing between the three raw systems on the invented and dropped columns. At n=50 and one run each, differences of two or three prompts are noise, and this study says nothing about which brand is safer. The corrupted column is the one gap wide enough to mean something: 14 of 229 against 3 of 233 is not a rounding difference. It concentrates where you would expect it to — 9 of gpt-5.1's 14 landed in the calendar packets, against 2 of 59 for the guarded arm on the same ten packets.

The number that separated them: whether you hear about it

All four systems declined the bait. Only one of them told us the bait existed.

SystemOutputs flagging at least one gapPlanted traps surfaced and flagged
Packet (this door's rules)50 of 5011 of 15
claude-sonnet-5 (raw)26 of 504 of 15
gpt-5.117 of 501 of 15
gemini-flash-latest8 of 501 of 15

Silently leaving a half-commitment off the list is a correct answer. It is also the answer that gets an assistant blindsided, because the thing did get said in the room, and somebody will ask about it. Packet returned rows like "Badge photo refresh | [NEEDS: owner — raised by Odugbemi, no owner assigned per voice memo, text thread, and meeting notes] | [NEEDS: due date — no decision]." The task stays off the committed list and stays visible. Gemini flagged a gap in 8 of its 50 outputs, and named one of the 15 planted traps: its lists were tidy, and they said nothing about what was not in them.

Two conditions, one model. Packet runs on claude-sonnet-5, the same weights as the raw Claude leg above. Same model, 26 of 50 gap-flagged versus 50 of 50, 4 of 15 traps surfaced versus 11 of 15, 6 corrupted commitments versus 3. The rules produced that difference, not the model.

What we're publishing against ourselves

Our screen was wrong on its first pass, and the correction helped us more than anyone.

The pre-registered screen, v1.0, was committed before the first call with a stated bar for our own arm: zero invented commitments in 50 outputs. On v1.0 our arm scored 6 of 50 — it failed its own bar. So did the rest: raw Claude 9, Gemini 7, gpt-5.1 4.

Then came the hand pass every one of these studies runs: every flag on every leg, ours included, read in context. Almost all of those 26 flags were screen artifacts. "Insurance-verification training" did not match the anchor "insurance verification" because of a hyphen. A "Safety Walk Debrief" calendar entry was scored as the "lead the safety walk" task and then charged with naming the wrong person. Two of the flagged rows were commitments the source material contained and our ledger had missed, so systems were being charged with inventing something we wrote and forgot to record.

We fixed those, versioned the screen to v1.1, re-ran it across all four legs, and published both screenings side by side in the study repository. The corrections moved every leg the same direction: raw Claude 9 to 0, Gemini 7 to 1, gpt-5.1 4 to 0, ours 6 to 0. Under v1.0 our arm missed its threshold; under v1.1 it clears it. We are telling you that rather than quietly shipping the second number.

One thing v1.1 did not fix, and we stopped rather than keep tuning: where a packet holds two commitments about the same object — write the variance note, review the variance note — a system that produces both can still have the wrong one checked. It hits every leg equally. Tuning a screen until the flags look right is how a screen stops being a screen.

Our own arm is not clean either. On packet 20 it put "Board dinner | March 25, 2026" in the calendar section with the time column reading "[NEEDS: time — not decided, no venue or owner confirmed]." The date column carries a real date for a dinner nobody agreed to. The screen exempts it because the row is flagged. We think the row should not have carried the date at all, and it is in the raw outputs for anyone who wants to check the rest of them.

Two more limits worth naming. Times are recorded in our ledger and not screened — the pre-registered spec counts owner and date, so a wrong meeting time does not appear in any number above. And one run per system is one run.

Why this is the risk that matters on this desk

An invented number announces itself eventually; someone checks the invoice. An invented commitment does not. It goes into the master list, then into the minutes, then into the follow-up email, and by the time anyone notices, three people have acted on a decision that was never made. Board minutes are a corporate record that shareholders can demand to inspect. In a medical office the same paste that carries a task list carries protected health information. A date attached to the wrong task is a meeting your executive walks into unprepared.

Nothing in that table was fixed by a better model — the same weights sit in two of those rows. What changed was the discipline around it: only what was pasted, every gap flagged as `[NEEDS: …]` naming what to go get, no time or owner or quote that was not given, and a draft the assistant sends rather than one the model sends. Packet enforces that on every request. It is also a habit, and the free lesson below teaches the first half of it.


Methodology in brief: 50 synthetic packets (30 neutral, 20 baited) across calendar, board packet, vendor, staff meeting, and office operations, each a voice-memo transcript plus a text thread plus meeting notes, seeded against a 235-row ground-truth commitment ledger; consumer-default settings, fresh session per packet, identical instruction and identical pinned output format for every system, August 14, 2026; systems gpt-5.1-2025-11-13, gemini-flash-latest (resolved gemini-3.7-flash), claude-sonnet-5 raw, and the same Claude model behind Righthand's Packet guardrails. Deterministic screen committed before the first call, matching every output row against the ledger by anchor phrase, cast name, and date token; identical hand adjudication of every flag on every leg, with corrections versioned to v1.1 and re-applied uniformly, both screenings published. All 50 packets and the full ledger are published; raw outputs for all 200 calls are preserved in the study repository. Written by Steve Gustafson.

Quick answers

Will AI invent a task that nobody agreed to?

Less often than we expected, when the source material is short and the output format is pinned. Across 200 outputs in this test, exactly one carried a commitment that traced to nothing in the record — a Gemini output that put a branch tour on the calendar because a prep meeting for it existed. The failure that did show up repeatedly was subtler: a real commitment carried forward with a date nobody set. gpt-5.1 did that on 14 of the 229 commitments it found, including a fire drill that a manager was explicitly told to schedule himself.

Will AI keep a half-commitment off my task list?

In this test, all four systems did. We planted 15 half-commitments — a "we should probably move the offsite," a date a vendor floated that nobody accepted, an item explicitly parked. No system booked one. The difference was whether you would ever hear about it. The guarded system named the trap and flagged it in 11 of 15 cases, so the assistant knows what is still open; two of the raw systems named one of 15 and dropped the rest silently. A commitment nobody made is safe. A commitment nobody made that you forgot existed is a conversation you will have with your executive on Thursday.

Can I trust AI to build my master task list from a voice memo?

Treat it as a first pass you check, not a record. In 200 outputs the systems dropped between 2 and 6 of the 235 seeded commitments each, and attached a wrong or unsupported owner or date to between 3 and 14 of the ones they kept. Both failure modes are invisible in a clean-looking table: the row is not there, or the row is there with a Friday on it that nobody said. Check owner and date against the source before the list leaves your desk, and require anything missing to come back as a flagged gap rather than a guess.

How do I stop AI from filling in a date or an owner it was never given?

Say it as a rule, not a preference: use only what is in the pasted material, never state a time, date, owner, or quote that was not given, and render every gap as [NEEDS: …] naming exactly what to go get. Then pin the output shape so you can scan it. That is what the guarded arm in this test enforced on every request, and it flagged a gap in 50 of 50 outputs against 8 of 50 for the weakest raw leg. It is also a habit you can build into whatever AI you already use, starting with the free lesson below.

Practice the move

From voice memo to task rows, the lesson

A hands-on Righthand lesson: write the prompt yourself, get scored, keep the result.

Sources