← Back to portfolio
An aircraft mechanic in a dark hangar reading an open maintenance manual in front of an opened jet engine

The weakest part of my AI pipeline was me

1 September 2026 · AI engineering · Human-in-the-loop · Plain language

325
bytes in the tool
15×
times invoked
1–2 min
re-pitch to decision
2
logged disagreements

1. The decision I could not answer

This landed in my terminal one morning:

Option A (triage recommends): add both. One new FR — "the assistant is structurally scoped to the band business and refuses off-topic requests" — plus an FR-21-style adversarial acceptance test. The LLM engineer's eval set gets off-topic probes.

I read it three times. Every word was correct. The recommendation was sound. I know that now. But sitting there, I could not have explained the decision back to anyone, and an answer you cannot explain back is a guess.

Two things were happening at once. The first is general: language models default to dense, formal prose. They mirror the register of what they reason over, so a finding about platform policy comes out sounding like platform policy, and a finding about system design comes out sounding like a design document. The second is personal: I am not a native English speaker, which makes that wall higher. But the wall is there for everyone. "Structurally scoped" and "adversarial acceptance test" stop a native speaker who is not deep in this project just as well.

2. A decision gate only works if the human understands the decision

Context, briefly. I am building a WhatsApp assistant for a wedding and event band business, and AI agents write the specifications. A separate audit agent, the reconciler, reads the documents against each other and reports every contradiction and gap. I wrote about that pipeline in the previous case study.

The pipeline's whole design rests on one rule: the agents write, the human decides. Every audit finding becomes one decision: options, trade-offs, a recommendation. I answer it. That is the human gate, and it is the part everyone praises when they say "human-in-the-loop".

Here is the quiet failure mode nobody instruments: the gate assumes the human understands the question. When the human doesn't, the gate still looks like it works. Decisions still get made. But an approval you do not understand is not an approval. It is a rubber stamp with extra steps.

3. The whole tool fits in this figure, and I didn't write it

~/.claude/skills/wait-what/SKILL.md, the complete file, 325 bytes
---
name: wait-what
description: Stop. That last message did not land — re-pitch it.
disable-model-invocation: true
---
 
Wait — I don't understand where you've got to here. Re-pitch that:
give me a little bit of context, talk in ASD-STE100 Simplified
Technical English, and use the ubiquitous language from `CONTEXT.md`.
The whole tool. One paragraph, no code, no framework. Created by Matt Pocock; I installed it as a personal skill and pointed it at my project's vocabulary.

The skill is Matt Pocock's idea, not mine. It is one instruction: stop, that did not land, pitch it again. Two details in it do the heavy lifting.

ASD-STE100 is Simplified Technical English, a controlled-language standard written for aircraft maintenance manuals. Its rules are blunt: short sentences, one instruction per sentence, a limited word list, the same word for the same thing every time. It exists because a mechanic misreading a manual is how planes fall out of the sky. It turns out to be exactly the register you want when a tired human must make a decision they will own.

disable-model-invocation: true means the model can never trigger this skill by itself. Only I can, by typing /wait-what. That is the right boundary, and it is worth stating plainly: the AI cannot detect that you didn't understand something. Confusion is private. The escape hatch must sit on the human side of the loop, because only the human knows when the message did not land.

4. Before and after, twice

What the re-pitch actually does is easiest to show. Both examples are quoted from my session transcripts, trimmed for length and nothing else.

Example one: a required behavior with no mechanism. The audit found that a P0 requirement existed in one spec but no component implemented it:

before: the audit presents finding 2
The conversation spec assigns the job to code: "a worker timer
registers the partial capture state when a thread goes quiet
mid-capture". The architecture spec never received the assignment.
Its list of worker jobs has six entries — export, audit, retention,
renewal, urgency, heartbeat — and no stall timer.
 
So a P0 behavior has a name but no mechanism anywhere. A or B?
Correct, complete, and a wall. I typed /wait-what finding 2.
after: the same decision, re-pitched
A client writes on WhatsApp. The assistant asks questions and
collects the event details: the date, the event type, the city.
This collection is called capture.
 
When enough details are known, the assistant saves them as an event
request in the database. This save action is called registration.
 
A client who stops answering is a problem for the assistant. The
assistant only acts when a message comes in. No message comes in, so
the assistant cannot act. Only a worker can act. A worker is a
program that runs in the background, on a clock, without messages.
 
Option A — register early. The assistant registers the event request
at the first moment the minimum is known. A stalled thread then
needs no timer at all. It removes a machine instead of adding one.
Same facts, same options. I answered within two minutes.

Example two: one undefined word. Mid-decision, a finding used a term as if everyone knew it: "the echo-storm backstop — if echo filtering ever breaks, the bot answers itself in a loop." I stopped and typed /wait-what what is echo filter?. The answer opened like this:

Under Coexistence, one phone number has two mouths. Mouth one: the manager's own WhatsApp app. Mouth two: our system, through Meta's API. Our system listens to Meta through a webhook — a door where Meta delivers every event about the number. When the manager sends a message from his phone, Meta delivers a copy of that message to our webhook too. That copy is called an echo. The echo filter is the code at the door that sorts them.

One number, two mouths, a door, and the copies that come back through it. After that paragraph, the original decision took under five minutes.

5. What every re-pitch did

Fifteen invocations later, the pattern is stable enough to write down as a recipe. Every good re-pitch made the same five moves:

  1. It opened with a concrete scene, with clock times. "A client writes on Monday, 10:00. This opens the service window. The window closes Tuesday, 10:00." Abstractions came only after the scene existed to hang them on.
  2. It defined every term before using it. Often as an explicit block: "Some terms first." A definition after first use arrives too late. The reader already stumbled.
  3. It gave the abstract thing one physical picture. A number with two mouths. A webhook as a door. A retention rule for records that nothing creates: "a shelf with no jars."
  4. It kept the same option letters. A stayed A and B stayed B, so the re-pitch never forced me to re-map the decision I was already holding.
  5. It kept the recommendation but lowered its voice. The advice survived the translation; the pressure did not.

None of these moves is clever. That is the point. They are the moves a good senior engineer makes when explaining a trade-off to a client, which is what this whole exercise turns out to be, with the roles reversed.

6. The numbers, and the real test of an explanation

The counts, from the session transcripts: I invoked /wait-what 15 times across two audit waves: 13 times in the two days of the first wave's triage, and twice, so far, in the second wave's. One invocation I aborted after four seconds and re-typed, so 14 re-pitches were actually delivered.

Ten of the fourteen ended in a decision within minutes, the fast ones in one to two. But the interesting rows are the other four, because none of them was a failure. Each one was the re-pitch converting confusion into a better question:

A good explanation does not make you agree. It makes you able to disagree, and then able to ask the next question. Twice in fifteen uses, the plain-language version changed the outcome away from the AI's own recommendation. Those two decisions are the strongest evidence in this post that the gate was previously letting things through it should not have.

7. The escape hatch became the default

After two days and thirteen invocations, I said the quiet part to the agent in plain words: I am not a native English speaker, and the complicated phrasing in the documents was hard for me to read and to understand, and the /wait-what output was not.

An escape hatch that fires thirteen times in two days is not a workaround. It is data about the default. So the default changed: simple English is now the written contract for every agent on the project. Each agent definition carries the rule: short sentences, terms defined at first use, the same word for the same thing. The three Wave 2 specifications were rewritten under it.

In the second wave's audit round, run against documents written under the new contract, I have needed /wait-what twice so far, down from thirteen.

This post is written in that style too. You have been reading the method the whole time.

8. What this is really about

It is tempting to file this under accessibility: a tool for non-native speakers. That undersells it. The human gate is a component, like the webhook or the database. It has an input format, a throughput, and a failure mode. And its failure mode is silent: decisions keep flowing whether or not the human understood them. I would never ship a parser that silently accepts input it cannot parse. I was shipping myself as exactly that.

The fix was not to make the human smarter or the model smarter. It was 325 bytes of interface change, written by someone else, plus the honesty to type it when needed. The measurable result: faster decisions, two logged disagreements with the AI's recommendations, one claim challenged at the source, and a default that improved for every future document.

There is also a second reading, and it's the one I'd put in a job interview. Explaining a technical trade-off in the listener's language, until the listener can push back, is not a nice-to-have communication skill. It is the job: with a non-technical client, with a stakeholder, and, it turns out, with yourself whenever you are the one at the gate. The band manager this product serves will never read a specification. Every decision he owns will have to survive exactly this translation. Now I know, from the inside, what that translation must do.


Fact table: every number in this post, and where it comes from

Counts marked "session transcripts" come from my private Claude Code session logs, which are not in the public repo. The grep is shown so the method is plain, and all counts are as of 2026-08-30. The second audit wave was still running when this post was written.

ClaimSource
325 byteswc -c ~/.claude/skills/wait-what/SKILL.md
Skill created by Matt Pocockhis published prompt; installed by me as a personal skill
15 invocationssession transcripts — grep -c for the /wait-what command marker across both audit sessions
13 in two dayssession transcripts — invocation timestamps: 8 on 2026-08-26, 5 on 2026-08-27
14 delivered re-pitches15 minus one invocation aborted four seconds in (2026-08-26 11:35) and immediately re-typed
10 of 14 decided within minutessession transcripts — re-pitch timestamp to my reply, per invocation
2 logged disagreementsband-platform docs/decisions.md — "FR-24 loses its client-facing courtesy message" (chose C against recommended A); glossary entry ruling out the proposed term after "give me more options"
Finding 2 quotes (before/after)session transcript, 2026-08-30, quoted verbatim and trimmed
Echo-filter quotessession transcript, 2026-08-27, quoted verbatim and trimmed
"are you sure… where did you read it?"session transcript, 2026-08-26, my reply after the T1.4 re-pitch, verbatim including my typo
Saturday-night escalation scenesession transcript, 2026-08-30, the finding-9 follow-up, quoted verbatim and trimmed
Simple English as agent contractband-platform .claude/agents/*.md — the "Writing style — simple English" section in each agent definition
Wave-2 usage: 2 so farsession transcripts — same grep, second session file; re-checked on publish day
ASD-STE100the ASD Simplified Technical English specification, asd-ste100.org
Hero imageAI-generated from a text prompt — no real hangar was photographed

Building something like this?

Open to senior frontend / full-stack roles — remote or Timișoara.

Get in touch