
The weakest part of my AI pipeline was me
1 September 2026 · AI engineering · Human-in-the-loop · Plain language
- 325
- bytes in the tool
- 15×
- times invoked
- 1–2 min
- re-pitch to decision
- 2
- logged disagreements
1. The decision I could not answer
This landed in my terminal one morning:
Option A (triage recommends): add both. One new FR — "the assistant is structurally scoped to the band business and refuses off-topic requests" — plus an FR-21-style adversarial acceptance test. The LLM engineer's eval set gets off-topic probes.
I read it three times. Every word was correct. The recommendation was sound. I know that now. But sitting there, I could not have explained the decision back to anyone, and an answer you cannot explain back is a guess.
Two things were happening at once. The first is general: language models default to dense, formal prose. They mirror the register of what they reason over, so a finding about platform policy comes out sounding like platform policy, and a finding about system design comes out sounding like a design document. The second is personal: I am not a native English speaker, which makes that wall higher. But the wall is there for everyone. "Structurally scoped" and "adversarial acceptance test" stop a native speaker who is not deep in this project just as well.
2. A decision gate only works if the human understands the decision
Context, briefly. I am building a WhatsApp assistant for a wedding and event band business, and AI agents write the specifications. A separate audit agent, the reconciler, reads the documents against each other and reports every contradiction and gap. I wrote about that pipeline in the previous case study.
The pipeline's whole design rests on one rule: the agents write, the human decides. Every audit finding becomes one decision: options, trade-offs, a recommendation. I answer it. That is the human gate, and it is the part everyone praises when they say "human-in-the-loop".
Here is the quiet failure mode nobody instruments: the gate assumes the human understands the question. When the human doesn't, the gate still looks like it works. Decisions still get made. But an approval you do not understand is not an approval. It is a rubber stamp with extra steps.
3. The whole tool fits in this figure, and I didn't write it
~/.claude/skills/wait-what/SKILL.md, the complete file, 325 bytes---name: wait-whatdescription: Stop. That last message did not land — re-pitch it.disable-model-invocation: true---Wait — I don't understand where you've got to here. Re-pitch that:give me a little bit of context, talk in ASD-STE100 SimplifiedTechnical English, and use the ubiquitous language from `CONTEXT.md`.
The skill is Matt Pocock's idea, not mine. It is one instruction: stop, that did not land, pitch it again. Two details in it do the heavy lifting.
ASD-STE100 is Simplified Technical English, a controlled-language standard written for aircraft maintenance manuals. Its rules are blunt: short sentences, one instruction per sentence, a limited word list, the same word for the same thing every time. It exists because a mechanic misreading a manual is how planes fall out of the sky. It turns out to be exactly the register you want when a tired human must make a decision they will own.
disable-model-invocation: true means the model can never trigger this
skill by itself. Only I can, by typing /wait-what. That is the right
boundary, and it is worth stating plainly: the AI cannot detect that you didn't
understand something. Confusion is private. The escape hatch must sit on
the human side of the loop, because only the human knows when the message did
not land.
4. Before and after, twice
What the re-pitch actually does is easiest to show. Both examples are quoted from my session transcripts, trimmed for length and nothing else.
Example one: a required behavior with no mechanism. The audit found that a P0 requirement existed in one spec but no component implemented it:
before: the audit presents finding 2The conversation spec assigns the job to code: "a worker timerregisters the partial capture state when a thread goes quietmid-capture". The architecture spec never received the assignment.Its list of worker jobs has six entries — export, audit, retention,renewal, urgency, heartbeat — and no stall timer.So a P0 behavior has a name but no mechanism anywhere. A or B?
after: the same decision, re-pitchedA client writes on WhatsApp. The assistant asks questions andcollects the event details: the date, the event type, the city.This collection is called capture.When enough details are known, the assistant saves them as an eventrequest in the database. This save action is called registration.A client who stops answering is a problem for the assistant. Theassistant only acts when a message comes in. No message comes in, sothe assistant cannot act. Only a worker can act. A worker is aprogram that runs in the background, on a clock, without messages.Option A — register early. The assistant registers the event requestat the first moment the minimum is known. A stalled thread thenneeds no timer at all. It removes a machine instead of adding one.
Example two: one undefined word. Mid-decision, a finding used a term as
if everyone knew it: "the echo-storm backstop — if echo filtering ever
breaks, the bot answers itself in a loop." I stopped and typed
/wait-what what is echo filter?. The answer opened like this:
Under Coexistence, one phone number has two mouths. Mouth one: the manager's own WhatsApp app. Mouth two: our system, through Meta's API. Our system listens to Meta through a webhook — a door where Meta delivers every event about the number. When the manager sends a message from his phone, Meta delivers a copy of that message to our webhook too. That copy is called an echo. The echo filter is the code at the door that sorts them.
One number, two mouths, a door, and the copies that come back through it. After that paragraph, the original decision took under five minutes.
5. What every re-pitch did
Fifteen invocations later, the pattern is stable enough to write down as a recipe. Every good re-pitch made the same five moves:
- It opened with a concrete scene, with clock times. "A client writes on Monday, 10:00. This opens the service window. The window closes Tuesday, 10:00." Abstractions came only after the scene existed to hang them on.
- It defined every term before using it. Often as an explicit block: "Some terms first." A definition after first use arrives too late. The reader already stumbled.
- It gave the abstract thing one physical picture. A number with two mouths. A webhook as a door. A retention rule for records that nothing creates: "a shelf with no jars."
- It kept the same option letters. A stayed A and B stayed B, so the re-pitch never forced me to re-map the decision I was already holding.
- It kept the recommendation but lowered its voice. The advice survived the translation; the pressure did not.
None of these moves is clever. That is the point. They are the moves a good senior engineer makes when explaining a trade-off to a client, which is what this whole exercise turns out to be, with the roles reversed.
6. The numbers, and the real test of an explanation
The counts, from the session transcripts: I invoked /wait-what 15 times
across two audit waves: 13 times in the two days of the first wave's triage,
and twice, so far, in the second wave's. One invocation I aborted after four
seconds and re-typed, so 14 re-pitches were actually delivered.
Ten of the fourteen ended in a decision within minutes, the fast ones in one to two. But the interesting rows are the other four, because none of them was a failure. Each one was the re-pitch converting confusion into a better question:
- Once I disagreed with the recommendation outright. The audit recommended deferring a classification question on a minor planned feature. The plain version laid the consequence chain bare, and a third option became obvious: cut that part of the feature, and the whole question disappears. I chose the cut. It is in the decision log.
- Once I sent the re-pitch back for alternatives. The plain version of a
naming decision let me see what the proposed term would cost us later. "Give
me more options." And the project's core domain term,
event_request, came out of that second round instead of the AI's first proposal. - Once I challenged the claims themselves. A re-pitch listed the permanent side effects the product would impose on my client's personal phone. My reply: "are you sure about all this limitations? can you show me where did you read it?" The plain version had made the claims concrete enough to doubt. Jargon never gets that far.
- Once I asked the next question down. See below.
A good explanation does not make you agree. It makes you able to disagree, and then able to ask the next question. Twice in fifteen uses, the plain-language version changed the outcome away from the AI's own recommendation. Those two decisions are the strongest evidence in this post that the gate was previously letting things through it should not have.
7. The escape hatch became the default
After two days and thirteen invocations, I said the quiet part to the agent
in plain words: I am not a native English speaker, and the complicated
phrasing in the documents was hard for me to read and to understand, and
the /wait-what output was not.
An escape hatch that fires thirteen times in two days is not a workaround. It is data about the default. So the default changed: simple English is now the written contract for every agent on the project. Each agent definition carries the rule: short sentences, terms defined at first use, the same word for the same thing. The three Wave 2 specifications were rewritten under it.
In the second wave's audit round, run against documents written under the new
contract, I have needed /wait-what twice so far, down from thirteen.
This post is written in that style too. You have been reading the method the whole time.
8. What this is really about
It is tempting to file this under accessibility: a tool for non-native speakers. That undersells it. The human gate is a component, like the webhook or the database. It has an input format, a throughput, and a failure mode. And its failure mode is silent: decisions keep flowing whether or not the human understood them. I would never ship a parser that silently accepts input it cannot parse. I was shipping myself as exactly that.
The fix was not to make the human smarter or the model smarter. It was 325 bytes of interface change, written by someone else, plus the honesty to type it when needed. The measurable result: faster decisions, two logged disagreements with the AI's recommendations, one claim challenged at the source, and a default that improved for every future document.
There is also a second reading, and it's the one I'd put in a job interview. Explaining a technical trade-off in the listener's language, until the listener can push back, is not a nice-to-have communication skill. It is the job: with a non-technical client, with a stakeholder, and, it turns out, with yourself whenever you are the one at the gate. The band manager this product serves will never read a specification. Every decision he owns will have to survive exactly this translation. Now I know, from the inside, what that translation must do.
Fact table: every number in this post, and where it comes from
Counts marked "session transcripts" come from my private Claude Code session logs, which are not in the public repo. The grep is shown so the method is plain, and all counts are as of 2026-08-30. The second audit wave was still running when this post was written.
| Claim | Source |
|---|---|
| 325 bytes | wc -c ~/.claude/skills/wait-what/SKILL.md |
| Skill created by Matt Pocock | his published prompt; installed by me as a personal skill |
| 15 invocations | session transcripts — grep -c for the /wait-what command marker across both audit sessions |
| 13 in two days | session transcripts — invocation timestamps: 8 on 2026-08-26, 5 on 2026-08-27 |
| 14 delivered re-pitches | 15 minus one invocation aborted four seconds in (2026-08-26 11:35) and immediately re-typed |
| 10 of 14 decided within minutes | session transcripts — re-pitch timestamp to my reply, per invocation |
| 2 logged disagreements | band-platform docs/decisions.md — "FR-24 loses its client-facing courtesy message" (chose C against recommended A); glossary entry ruling out the proposed term after "give me more options" |
| Finding 2 quotes (before/after) | session transcript, 2026-08-30, quoted verbatim and trimmed |
| Echo-filter quotes | session transcript, 2026-08-27, quoted verbatim and trimmed |
| "are you sure… where did you read it?" | session transcript, 2026-08-26, my reply after the T1.4 re-pitch, verbatim including my typo |
| Saturday-night escalation scene | session transcript, 2026-08-30, the finding-9 follow-up, quoted verbatim and trimmed |
| Simple English as agent contract | band-platform .claude/agents/*.md — the "Writing style — simple English" section in each agent definition |
| Wave-2 usage: 2 so far | session transcripts — same grep, second session file; re-checked on publish day |
| ASD-STE100 | the ASD Simplified Technical English specification, asd-ste100.org |
| Hero image | AI-generated from a text prompt — no real hangar was photographed |