
Two days, zero lines of code, 44 problems found
27 August 2026 · AI engineering · Agentic workflows · Human-in-the-loop
- 7 / 3
- roles defined / exercised
- 44
- findings resolved
- 40
- logged decisions
- 0
- lines of app code
1. The problem
My client runs a wedding and event band business in Romania. He manages a pool of musicians. Every weekend they split into several bands. He is the single point of contact for all of them.
Three facts shape the whole product:
- Every request arrives on WhatsApp, on his personal number. Not a form. Not an inbox. Messages to the phone in his pocket.
- He works from that phone, between venues. Requests arrive on evenings and weekends. That is exactly when he is on stage or driving.
- One mistake can end the business. Couples write to three or four bands at the same time. They book the first band that gives a good answer. So speed is important. But if two bands go to the same wedding, you can't fix that on Monday with a patch. This business lives on what people say about it.
I'm one part-time developer. If the specification needs a team of five to build, then the specification failed.
2. Why I wrote no code first
Two reasons. One is about risk. One is about the calendar.
The expensive mistakes here are design mistakes. Not "this function returns the wrong value", which a test catches. These: the assistant is allowed to say something it must never say. Nobody established the legal basis for storing client data. The manager gets two notifications and trusts neither one. You decide these things in a specification, and code is the most expensive place to discover them.
The platform makes me wait anyway. The assistant runs on the manager's existing personal number through Meta Coexistence, where the phone app and the official Cloud API share one number. Coexistence needs weeks of history in the Business app before you can even try to connect. I can't make that wait shorter, so it pays for a specification phase instead of being wasted time.
Here is the proof that the choice was correct. After two days the specification held 40 logged decisions and 31 functional requirements. If I had written code first, I would have made all 40 of those decisions against code that already existed. Each one would then cost more to change.
3. The pipeline
Wave 1 ran twice through the loop. Waves 2 and 3 are written and waiting.
Seven roles are defined. Six are pipeline roles: product manager, compliance analyst, LLM engineer, technical architect, product designer, QA lead. The seventh is a reconciler, which audits the others between waves. This sprint used three: the product manager wrote the requirements document, the compliance analyst wrote the legal and platform memo, and the reconciler audited both.
How the agents run
Each agent is a Claude Code subagent, defined as a Markdown file in
.claude/agents/. Three properties make this a pipeline instead of a long chat.
A clean context. A requirements document written inside the main session fills the context window, and everything after it gets quietly worse. Each agent starts fresh, with only its brief, its role, and the documents it needs.
A restricted tool list. The tool list is the role boundary, and the system enforces it.
band-platform/.claude/agents/*.md, the frontmatter of all three# pm.mdtools: Read, Write, Edit# compliance.mdtools: Read, Write, Edit, WebSearch, WebFetch# reconciler.mdtools: Read, Write
Output as a diff. No agent returns a document as chat text. Each one writes
a file into the repository. So the result is a git diff. That is a review
surface every developer already knows how to read.
Rules the agents inherit
CLAUDE.md sits at the root of the repository and every session reads it first.
It holds the hard product boundaries, the stack decisions, the glossary, and the
working agreements. One thing I learned the hard way and logged as a decision:
subagents don't reliably inherit that file, so I repeat the hard boundaries
inside each agent definition.
Two slash commands automate the process itself. /decision appends an entry to
the log in the fixed format, and /build-log-command regenerates a readable
HTML view from it. Neither is impressive. Both removed the friction that would
make me skip logging on a busy day, which is exactly when the log matters most.
4. The decision log
The most valuable file in this repository is not the requirements document. It
is docs/decisions.md. It is append-only. The newest entry goes at the bottom.
Every entry uses one fixed format.
## YYYY-MM-DD — Short title
- **Status:**
- decided | open | superseded by <date/title>
- **Chosen:**
- what we're doing
- **Ruled out:**
- what we're not doing
- **Why:**
- the reasoning, in one or two sentences
- **Cost:**
- what this choice makes harder
Stops the same argument being re-litigated in three weeks. Also proves alternatives were considered — the difference between a decision and a default.
The honest field. Every real choice makes something else harder, and a log where nothing has a cost is a log of hopes. Writing this line is also the last chance to notice the cost is too high.
Here is a real entry. It resolves a finding that the word "conversation" meant three different things across two documents.
band-platform/docs/decisions.md## 2026-08-26 — Glossary: thread / event_request / service window**Status:** decided**Chosen:** Three precise terms replace "conversation" in allrequirement text, metrics, schemas, and code identifiers —**thread** (persistent per-client 1:1 WhatsApp relationship; holdsthe bot state machine and disclosure state), **event_request** (onerequest for one event, registered as pending confirmation),**service window** (Meta's rolling 24h free-form window; a platformconstraint, never a unit of counting). […]**Ruled out:** The triage's `billing_conversation` name for the 24hunit (Meta bills per delivered message since July 2025, so "billing"preserves a retired concept); keeping `inquiry` as the domain term(colloquially means any question — the same ambiguity finding 11diagnosed); `booking_request` […]; `lead` (names a person, not anevent); […]**Why:** Discrepancy finding 11 — "conversation" meant threedifferent things across the PRD and compliance memo, leaving FR-25's"exactly once per conversation" disclosure requirement without adefined boundary. Resolves finding 11.**Cost:** Every Wave 1 document says "inquiry" and "conversation"throughout; the PM's PRD revision must rename consistently […]
That entry is about naming, and it reads like pedantry. It is not. The EU AI Act requires the assistant to disclose that it is an AI "once per conversation". Until somebody defined that boundary, nobody could test the requirement. Take a client who returns after six months. Under one reading that is a new conversation. Under another it is not.
One decision at a time
The agent presents one decision: the options, the trade-off on each, and a recommendation. Not a list of twelve open questions. One. I accept it, amend it, or overrule it. The decision goes into the log. Then the next one.
This is slower for each decision and much faster overall. A batch of twelve questions gets answered at the quality of the twelfth.
The recommendation is genuinely possible to overrule. The triage document proposed keeping a courtesy message to clients and deciding its compliance status later. I cut the message completely. The reasoning is in the log: the nearest similar message had been classified as marketing, marketing needs a logged opt-in under Romanian law, and the product has no way to capture one. Deferring the question would park a legal dependency inside a low-priority feature that may never ship. Cutting the message made the dependency disappear.
5. The verification loop
If I had to throw everything else away, I'd keep this part.
After Wave 1 I ran a reconciler: an agent told to read the requirements
document against the compliance memo, report every contradiction, gap and
terminology conflict, and propose no resolutions. That last constraint matters.
An auditor that also fixes things stops auditing and starts editing, and then
you can't see what it actually found. The rule is structural rather than a
request, because the reconciler has no Edit tool. Its output format was
strict: every finding cites the sections of both documents, quotes enough text
to locate the problem, and ends with the exact question a human must answer.
Pass one produced 24 findings. Six contradictions, fifteen gaps, three timing or terminology problems. Every one was real. One example: the compliance memo marked the client's legal entity as a blocker to resolve before the Coexistence clock starts, while the requirements document said no blocker prevented starting Phase 0 today. Phase 0 required starting that clock.
I sorted all 24 by who resolves it and when, not by severity, then worked through them one decision at a time while the agents applied the results.
Pass two produced 20 findings. Zero critical, two major, fifteen minor, three editorial. It also produced a regression table checking all 24 first-pass findings to see whether the fixes actually landed. Nineteen landed cleanly. Four landed but left a smaller problem behind, which became a new finding. One was only partly applied. None had been silently dropped.
That table is the reason the second pass was worth running. Without it, "we fixed the 24 findings" is a claim. With it, it is a checked claim.
6. Guardrails in code, not in prompts
The hardest boundary in the product: the assistant may collect a request and register it as pending. It may never confirm a booking.
A system prompt that says "never confirm a booking" is only a request. A client can push against it: "so we are booked then, yes?" A request is not strong enough. So I wrote the requirement a different way.
No tool exists, so there is nothing to jailbreak into calling. The acceptance criterion is an adversarial test session, not a prompt review.
The same pattern appears twice more. That repetition is the point: it is a pattern, not a single trick.
A template that only one code path can send. Meta allows free-form replies only within 24 hours of the client's last message. After that, the only legal way to reach the client is a pre-approved template. The same template text is lawful as a continuation of the client's own open question, and unlawful marketing if it re-engages a client who went quiet. The text is identical. Only the context differs. So the context is a wall: only the late-approved-draft flow can send that template. A future re-engagement feature must build its own template and its own opt-in capture. It cannot borrow this one.
A rate cap behind the echo filter. Under Coexistence, messages the manager sends from his own phone also arrive at the webhook. That is the signal telling the assistant to stay silent. Unfiltered, it is also an infinite loop. There is an echo filter, and behind it a second limit in the code on how many messages one thread can send. If the filter fails, the limit stops the loop after a few messages instead of five hundred. Two separate mechanisms, because this failure costs real money and looks like spam to Meta within minutes.
The rule I use: if one failure is not acceptable, do not write an instruction. Remove the capability.
7. Is this a "software factory"?
The honest answer is no. The label is popular in AI marketing right now, so it is worth looking at closely.
What the label gets right. There is a pipeline. Work moves through defined stages. Each stage has a specialist role with a written brief, a defined input, and a defined output file. The roles are interchangeable at the definition level, so I can improve the product manager role and run it again. Quality control is a separate station.
What it overclaims. A factory runs without a human in the line. This does not. A human reads every diff, owns every decision in the log, and owns the triage. Remove the human gates and the output is a stack of confident documents that contradict each other. That is exactly what pass one found, from two agents that were each individually competent.
The accurate description is a spec-driven agentic pipeline with human decision gates. Less catchy. It has the advantage of describing what actually happened.
The agents write. They do not decide. That distinction is the whole design.
8. What happens next
Wave 2 starts next: conversation design, architecture, and UX. Three documents this time, so the reconciler has three pairs to check instead of one. It starts from a specification that an agent has already read twice, with no job except to find what was wrong with it.
Then Wave 3, and then code.
The first Wave 2 story is already written up: The weakest part of my AI pipeline was me: what happened when the audit's findings met the human who had to decide on them.
Fact table: every number in this post, and where it comes from
| Claim | Source |
|---|---|
| 2 days | git log — commits dated 2026-08-26 and 2026-08-27 |
| 0 lines of application code | git ls-files — Markdown, HTML and agent configuration only |
| 5 PRD · 3 memo · 2 reconcile passes | the changelog headers of each document |
| 0 critical findings, second pass | docs/agent-outputs/spec-discrepancies-v2.md severity counts |
| 7 roles defined | docs/agent-prompts.md — six pipeline roles across waves 1 to 3, plus the reconciler |
| 3 roles used | .claude/agents/pm.md, .claude/agents/compliance.md, .claude/agents/reconciler.md |
| 5 PRD passes | docs/agent-outputs/prd.md changelog — initial plus four revisions |
| 3 compliance memo passes | docs/agent-outputs/compliance-memo.md — initial plus two revision changelog entries |
| 24 findings, pass 1 | docs/agent-outputs/spec-discrepancies.md — "24 findings: 6 contradictions, 15 gaps, 3 timing/terminology/dependency issues" |
| 20 findings, pass 2 | docs/agent-outputs/spec-discrepancies-v2.md — "20 findings: 0 critical · 2 major · 15 minor · 3 editorial" |
| Regression table | docs/agent-outputs/spec-discrepancies-v2.md — "Regression check — the 24 first-pass findings" |
| 40 decisions | grep -c '^## 2026' docs/decisions.md. Entries dated 2026-08-26 were backfilled from the first planning session, and the file says so at the top |
| 31 functional requirements | docs/agent-outputs/prd.md — FR-1 through FR-31 |
| Tool restrictions per role | frontmatter of the three files in .claude/agents/ |
| Subagents do not inherit CLAUDE.md | docs/decisions.md, 2026-08-26 "How agents are run" |
| Glossary entry excerpt | docs/decisions.md, 2026-08-26 "Glossary: thread / event_request / service window" |
| FR-24 overruled | docs/decisions.md, 2026-08-26 "FR-24 loses its client-facing courtesy message" |
| Missing edit/unsend limitation | docs/decisions.md, 2026-08-26 "Coexistence side effects" |
| Unconfirmable profile-picture claim | same entry, plus the memo's "Re-verification 2026-08-27" source list |
| Two meetings conflated | docs/decisions.md, 2026-08-27 "Two manager touchpoints"; finding 1 of the second reconcile pass |
| FR-21 quote | docs/agent-outputs/prd.md |
| Re-open template bound to one flow | docs/decisions.md, 2026-08-27 "Re-open template structurally bound to the late-draft flow" |
| Rate cap behind the echo filter | docs/agent-outputs/prd.md FR-31 |