← Back to portfolio
A conductor on a podium directing an orchestra, seen from the back of a darkened concert hall

Two days, zero lines of code, 44 problems found

27 August 2026 · AI engineering · Agentic workflows · Human-in-the-loop

7 / 3
roles defined / exercised
44
findings resolved
40
logged decisions
0
lines of app code

1. The problem

My client runs a wedding and event band business in Romania. He manages a pool of musicians. Every weekend they split into several bands. He is the single point of contact for all of them.

Three facts shape the whole product:

I'm one part-time developer. If the specification needs a team of five to build, then the specification failed.

2. Why I wrote no code first

Two reasons. One is about risk. One is about the calendar.

The expensive mistakes here are design mistakes. Not "this function returns the wrong value", which a test catches. These: the assistant is allowed to say something it must never say. Nobody established the legal basis for storing client data. The manager gets two notifications and trusts neither one. You decide these things in a specification, and code is the most expensive place to discover them.

The platform makes me wait anyway. The assistant runs on the manager's existing personal number through Meta Coexistence, where the phone app and the official Cloud API share one number. Coexistence needs weeks of history in the Business app before you can even try to connect. I can't make that wait shorter, so it pays for a specification phase instead of being wasted time.

Here is the proof that the choice was correct. After two days the specification held 40 logged decisions and 31 functional requirements. If I had written code first, I would have made all 40 of those decisions against code that already existed. Each one would then cost more to change.

3. The pipeline

The agent pipeline: waves of specialist agents with human gates between themA project brief feeds Wave 1, where a product manager agent and a compliance analyst agent write in parallel. A human gate follows. A reconciler agent audits both documents. A human triages the findings and logs decisions, which loop back into a revision of Wave 1. Wave 2 and Wave 3 have not run yet.Project brief (human-written)Product manager agentwrites prd.mdCompliance analyst agentwrites compliance-memo.mdHUMAN GATE — read every diffReconciler agentaudits PRD against memo, emits findingsHUMAN — triage, then decideevery decision appended to decisions.mdWave 2 — LLM engineer · architect · designerWave 3 — QA, then codeWave 1 — run, in parallel, isolated contextsrevise and re-runnot yet runnot yet run

Wave 1 ran twice through the loop. Waves 2 and 3 are written and waiting.

Seven roles are defined. Six are pipeline roles: product manager, compliance analyst, LLM engineer, technical architect, product designer, QA lead. The seventh is a reconciler, which audits the others between waves. This sprint used three: the product manager wrote the requirements document, the compliance analyst wrote the legal and platform memo, and the reconciler audited both.

How the agents run

Each agent is a Claude Code subagent, defined as a Markdown file in .claude/agents/. Three properties make this a pipeline instead of a long chat.

A clean context. A requirements document written inside the main session fills the context window, and everything after it gets quietly worse. Each agent starts fresh, with only its brief, its role, and the documents it needs.

A restricted tool list. The tool list is the role boundary, and the system enforces it.

band-platform/.claude/agents/*.md, the frontmatter of all three
# pm.md
tools: Read, Write, Edit
 
# compliance.md
tools: Read, Write, Edit, WebSearch, WebFetch
 
# reconciler.md
tools: Read, Write
Three roles, three different boundaries. The product manager has no web access, because a requirements document is written from the brief and not from the internet. The compliance analyst has web access, because it must check current platform policy and current law. That difference mattered a great deal, as section 5 shows. The reconciler has no Edit tool at all.

Output as a diff. No agent returns a document as chat text. Each one writes a file into the repository. So the result is a git diff. That is a review surface every developer already knows how to read.

Rules the agents inherit

CLAUDE.md sits at the root of the repository and every session reads it first. It holds the hard product boundaries, the stack decisions, the glossary, and the working agreements. One thing I learned the hard way and logged as a decision: subagents don't reliably inherit that file, so I repeat the hard boundaries inside each agent definition.

Two slash commands automate the process itself. /decision appends an entry to the log in the fixed format, and /build-log-command regenerates a readable HTML view from it. Neither is impressive. Both removed the friction that would make me skip logging on a busy day, which is exactly when the log matters most.

4. The decision log

The most valuable file in this repository is not the requirements document. It is docs/decisions.md. It is append-only. The newest entry goes at the bottom. Every entry uses one fixed format.

## YYYY-MM-DD — Short title

**Status:**
decided | open | superseded by <date/title>
**Chosen:**
what we're doing
**Ruled out:**
what we're not doing

Stops the same argument being re-litigated in three weeks. Also proves alternatives were considered — the difference between a decision and a default.

**Why:**
the reasoning, in one or two sentences
**Cost:**
what this choice makes harder

The honest field. Every real choice makes something else harder, and a log where nothing has a cost is a log of hopes. Writing this line is also the last chance to notice the cost is too high.

Six fields, fixed order, append-only. The two marked in orange are the ones that make the log worth keeping.

Here is a real entry. It resolves a finding that the word "conversation" meant three different things across two documents.

band-platform/docs/decisions.md
## 2026-08-26 — Glossary: thread / event_request / service window
 
**Status:** decided
**Chosen:** Three precise terms replace "conversation" in all
requirement text, metrics, schemas, and code identifiers —
**thread** (persistent per-client 1:1 WhatsApp relationship; holds
the bot state machine and disclosure state), **event_request** (one
request for one event, registered as pending confirmation),
**service window** (Meta's rolling 24h free-form window; a platform
constraint, never a unit of counting). […]
**Ruled out:** The triage's `billing_conversation` name for the 24h
unit (Meta bills per delivered message since July 2025, so "billing"
preserves a retired concept); keeping `inquiry` as the domain term
(colloquially means any question — the same ambiguity finding 11
diagnosed); `booking_request` […]; `lead` (names a person, not an
event); […]
**Why:** Discrepancy finding 11 — "conversation" meant three
different things across the PRD and compliance memo, leaving FR-25's
"exactly once per conversation" disclosure requirement without a
defined boundary. Resolves finding 11.
**Cost:** Every Wave 1 document says "inquiry" and "conversation"
throughout; the PM's PRD revision must rename consistently […]
Quoted verbatim. […] marks text cut for length; nothing else is changed.

That entry is about naming, and it reads like pedantry. It is not. The EU AI Act requires the assistant to disclose that it is an AI "once per conversation". Until somebody defined that boundary, nobody could test the requirement. Take a client who returns after six months. Under one reading that is a new conversation. Under another it is not.

One decision at a time

The agent presents one decision: the options, the trade-off on each, and a recommendation. Not a list of twelve open questions. One. I accept it, amend it, or overrule it. The decision goes into the log. Then the next one.

This is slower for each decision and much faster overall. A batch of twelve questions gets answered at the quality of the twelfth.

The recommendation is genuinely possible to overrule. The triage document proposed keeping a courtesy message to clients and deciding its compliance status later. I cut the message completely. The reasoning is in the log: the nearest similar message had been classified as marketing, marketing needs a logged opt-in under Romanian law, and the product has no way to capture one. Deferring the question would park a legal dependency inside a low-priority feature that may never ship. Cutting the message made the dependency disappear.

5. The verification loop

The verification loopDocuments are written, then reconciled by an agent, then triaged, then decided by a human, then applied back to the documents, which sends the cycle back to reconcile. Pass one produced 24 findings, pass two produced 20, with zero critical.1 · Writeagents draft2 · Reconcileadversarial agent3 · Triagesort by owner4 · Decidehuman, logged5 · Applyagents revise docsre-reconcilePass 1: 24 findings · Pass 2: 20 findings, 0 critical · 44 total, all resolvedBold outline = the step where the work is actually checked or owned

If I had to throw everything else away, I'd keep this part.

After Wave 1 I ran a reconciler: an agent told to read the requirements document against the compliance memo, report every contradiction, gap and terminology conflict, and propose no resolutions. That last constraint matters. An auditor that also fixes things stops auditing and starts editing, and then you can't see what it actually found. The rule is structural rather than a request, because the reconciler has no Edit tool. Its output format was strict: every finding cites the sections of both documents, quotes enough text to locate the problem, and ends with the exact question a human must answer.

Pass one produced 24 findings. Six contradictions, fifteen gaps, three timing or terminology problems. Every one was real. One example: the compliance memo marked the client's legal entity as a blocker to resolve before the Coexistence clock starts, while the requirements document said no blocker prevented starting Phase 0 today. Phase 0 required starting that clock.

How 44 reconciler findings were sorted into four tiers by ownerTwo reconcile passes produced 44 findings: 24 in the first pass and 20 in the second. They were sorted into four tiers by who resolves each one and when: blocking questions for the client that week, product decisions owned by the developer, legal obligations satisfied by a document rather than a feature, and recurring operational duties that had no owner. Every tier feeds the decision log.44 findings — 24 first pass, 20 secondHUMAN — sort by owner, not by severitywho resolves this, and when?TIER 1Blockingquestionsthe client, that weekTIER 2Productdecisionsmine to makeTIER 3Legalobligationsa document, not codeTIER 4Operationaldutiesnobody owned theseEvery outcome appended to decisions.mdthen the agents revise, and the reconciler runs againBold outline = the step a human owns

I sorted all 24 by who resolves it and when, not by severity, then worked through them one decision at a time while the agents applied the results.

Pass two produced 20 findings. Zero critical, two major, fifteen minor, three editorial. It also produced a regression table checking all 24 first-pass findings to see whether the fixes actually landed. Nineteen landed cleanly. Four landed but left a smaller problem behind, which became a new finding. One was only partly applied. None had been silently dropped.

That table is the reason the second pass was worth running. Without it, "we fixed the 24 findings" is a claim. With it, it is a checked claim.

6. Guardrails in code, not in prompts

The hardest boundary in the product: the assistant may collect a request and register it as pending. It may never confirm a booking.

A system prompt that says "never confirm a booking" is only a request. A client can push against it: "so we are booked then, yes?" A request is not strong enough. So I wrote the requirement a different way.

No tool exists, so there is nothing to jailbreak into calling. The acceptance criterion is an adversarial test session, not a prompt review.

The same pattern appears twice more. That repetition is the point: it is a pattern, not a single trick.

A template that only one code path can send. Meta allows free-form replies only within 24 hours of the client's last message. After that, the only legal way to reach the client is a pre-approved template. The same template text is lawful as a continuation of the client's own open question, and unlawful marketing if it re-engages a client who went quiet. The text is identical. Only the context differs. So the context is a wall: only the late-approved-draft flow can send that template. A future re-engagement feature must build its own template and its own opt-in capture. It cannot borrow this one.

A rate cap behind the echo filter. Under Coexistence, messages the manager sends from his own phone also arrive at the webhook. That is the signal telling the assistant to stay silent. Unfiltered, it is also an infinite loop. There is an echo filter, and behind it a second limit in the code on how many messages one thread can send. If the filter fails, the limit stops the loop after a few messages instead of five hundred. Two separate mechanisms, because this failure costs real money and looks like spam to Meta within minutes.

The rule I use: if one failure is not acceptable, do not write an instruction. Remove the capability.

7. Is this a "software factory"?

The honest answer is no. The label is popular in AI marketing right now, so it is worth looking at closely.

What the label gets right. There is a pipeline. Work moves through defined stages. Each stage has a specialist role with a written brief, a defined input, and a defined output file. The roles are interchangeable at the definition level, so I can improve the product manager role and run it again. Quality control is a separate station.

What it overclaims. A factory runs without a human in the line. This does not. A human reads every diff, owns every decision in the log, and owns the triage. Remove the human gates and the output is a stack of confident documents that contradict each other. That is exactly what pass one found, from two agents that were each individually competent.

The accurate description is a spec-driven agentic pipeline with human decision gates. Less catchy. It has the advantage of describing what actually happened.

The agents write. They do not decide. That distinction is the whole design.

8. What happens next

Wave 2 starts next: conversation design, architecture, and UX. Three documents this time, so the reconciler has three pairs to check instead of one. It starts from a specification that an agent has already read twice, with no job except to find what was wrong with it.

Then Wave 3, and then code.

The first Wave 2 story is already written up: The weakest part of my AI pipeline was me: what happened when the audit's findings met the human who had to decide on them.


Fact table: every number in this post, and where it comes from
ClaimSource
2 daysgit log — commits dated 2026-08-26 and 2026-08-27
0 lines of application codegit ls-files — Markdown, HTML and agent configuration only
5 PRD · 3 memo · 2 reconcile passesthe changelog headers of each document
0 critical findings, second passdocs/agent-outputs/spec-discrepancies-v2.md severity counts
7 roles defineddocs/agent-prompts.md — six pipeline roles across waves 1 to 3, plus the reconciler
3 roles used.claude/agents/pm.md, .claude/agents/compliance.md, .claude/agents/reconciler.md
5 PRD passesdocs/agent-outputs/prd.md changelog — initial plus four revisions
3 compliance memo passesdocs/agent-outputs/compliance-memo.md — initial plus two revision changelog entries
24 findings, pass 1docs/agent-outputs/spec-discrepancies.md — "24 findings: 6 contradictions, 15 gaps, 3 timing/terminology/dependency issues"
20 findings, pass 2docs/agent-outputs/spec-discrepancies-v2.md — "20 findings: 0 critical · 2 major · 15 minor · 3 editorial"
Regression tabledocs/agent-outputs/spec-discrepancies-v2.md — "Regression check — the 24 first-pass findings"
40 decisionsgrep -c '^## 2026' docs/decisions.md. Entries dated 2026-08-26 were backfilled from the first planning session, and the file says so at the top
31 functional requirementsdocs/agent-outputs/prd.md — FR-1 through FR-31
Tool restrictions per rolefrontmatter of the three files in .claude/agents/
Subagents do not inherit CLAUDE.mddocs/decisions.md, 2026-08-26 "How agents are run"
Glossary entry excerptdocs/decisions.md, 2026-08-26 "Glossary: thread / event_request / service window"
FR-24 overruleddocs/decisions.md, 2026-08-26 "FR-24 loses its client-facing courtesy message"
Missing edit/unsend limitationdocs/decisions.md, 2026-08-26 "Coexistence side effects"
Unconfirmable profile-picture claimsame entry, plus the memo's "Re-verification 2026-08-27" source list
Two meetings conflateddocs/decisions.md, 2026-08-27 "Two manager touchpoints"; finding 1 of the second reconcile pass
FR-21 quotedocs/agent-outputs/prd.md
Re-open template bound to one flowdocs/decisions.md, 2026-08-27 "Re-open template structurally bound to the late-draft flow"
Rate cap behind the echo filterdocs/agent-outputs/prd.md FR-31

Building something like this?

Open to senior frontend / full-stack roles — remote or Timișoara.

Get in touch