I decided to test different ways to “orchestrate agents” with some experiments. Then I was able to run uv run 01-orchestrator-worker/experiment.py --n 5 and, voilá! It went out to test different workers and trying to identify risks I planted in a codebase...
15 runs later, a clean-looking table with some discoveries and a potential inference. More workers/subagents are better, right? Yeah, sure... the version with the most workers found every planted risk. The versions with two workers missed one now and then.
Nice... except it was wrong.
One planted risk was an HTTP call with no timeout, written with httpx. And httpx has a 5-second timeout by default... so no risk at all. The two-worker versions that “missed” it were right, and the uncapped version got credit for a false alarm.
My agents were fine. My answer key was not... why? Because I wrote the key assuming no timeout and “forgot” that httpx has one by default...
I swapped it for a requests call (which really does wait forever) and re-ran all 15. First lesson of the series, before any architecture: whoever defines the right answer decides the result.
But Thiago, is this not just microservices again? A lead, some workers, a contract between them... Kinda :-P
I thought I was choosing an architecture, but I was drawing an org chart.
In Issue 12 I left you in the org chart and promised a terminal. Here it is... and it walked me straight back into the org chart. Its four rights (mandate, enforce, instruct, override) are about to become code.
Sooooo yeah, a series. 5 issues in 5 weeks, 1 public repo, 1 family of agent setups per week, each with a number from my own run. The repo is agent-topologies: seven folders planned, the first one live today, each README in plain language so you can follow it without writing a line of Python.
Ken Huang’s post “Multi-Agent Design Patterns” (Agentic AI Substack, August 2026) gave me food for thought. His catalogue sorts the patterns by shape. I think the question a CTO can actually act on is a different one: who owns the split?
If this lands, forward it to whoever is drawing your “agents’ org chart” :-D
Your agents will copy your org chart
Mel Conway wrote it down in “How Do Committees Invent?” (Datamation, April 1968): organisations that design systems “are constrained to produce designs which are copies of the communication structures of these organizations”.
And now 58y later, the communication structure has new members: Agents. Whoever decomposes the work for them decides the shape of what they produce.
And the org-design vocabulary transfers almost too well. Matthew Skelton and Manuel Pais’s Team Topologies (IT Revolution, 2019) names three ways teams should interact:
Collaboration: “working together for a defined period of time to discover new things”.
X-as-a-Service: “one team provides and one team consumes something ‘as a Service’”.
Facilitating: “temporary and focused”.
Bam!! Each one has an agent twin.
A worker with a strict input/output contract? X-as-a-Service.
A critic loop is collaboration, and it should be time-boxed... in agent land that is the iteration cap (coming with a code example soon).
A subagent that explores and hands an answer back is facilitating.
The neutral names for the patterns exist too. Erik S. and Barry Zhang’s “Building Effective Agents” (Anthropic, December 2024) names orchestrator-workers and evaluator-optimizer, and advises “the simplest solution possible, and only increasing complexity when needed”. The ownership question is what I add.
So the copy happens anyway. The only choice is whether you pick the original first.
The orchestrator is a team lead with a context window
Now the numbers from the repo, as promised.
Folder 01 is a risk review: eight short, fictional service descriptions (cards) with twelve risks planted in them, like a hardcoded key or an HTTP call with no timeout.
The lead reads a one-paragraph summary of each card and decides how many workers to use and which cards each gets. The workers review their cards and report in a fixed JSON format, never talking to each other. Then the lead merges everything into one summary for the CTO. A risk counts as found only with the right card and the right category.
Around them sit four safety checks, all in orchestrator.py:
Strict format (
FINDINGS_SCHEMA): one retry on a bad report, then it is thrown out and its cards listed as “not reviewed”. The X-as-a-Service contract, in code.Worker cap (
apply_fan_out_cap): too many workers get squeezed into fewer, no card dropped.Token budget (
dispatch): 60,000 tokens per review; once over, the remaining cards are skipped and named in the summary, not silently lost.Backup plan (
fallback_plan): if the lead fails, split the cards alphabetically. A failed lead never stops the review.
Three versions, five runs each (ran on 2026-Sep-25th): Claude Opus 5 as the lead, Claude Haiku 4.5 as the workers, through headless Claude Code (runs at commit 3a09186, after the answer-key fix. folder 01 is here).
Version Workers Tokens per review (average) Risks found Uncapped 4 36,573 60 of 60 Capped at two 2 23,278 59 of 60 No lead plan, alphabetical 2 22,994 59 of 60
Left alone, the lead planned four workers in all five runs. Capping it at two cut tokens by 36% for one missed risk out of 60. Skipping the lead’s plan and splitting the cards alphabetically between two workers: 37% fewer tokens, the same one miss.
So, Thiago, does the lead earn its keep? On this (simple) task, no. The lead’s planning call averaged 1,493 tokens per review. That is everything the plan cost, and it bought nothing I could measure over a dumb split with the same headcount.
More workers also meant more noise: 3.6 findings per review not on the answer key (the ones I spot-checked looked more like architectural calls the model was making for me than wrong answers), against 1.0 and 0.8 with two.
The caveats travel with the number:
Tokens, not cost. Cached input counts like fresh input, though it bills at about a tenth.
Per-call overhead. Claude Code adds a fixed prompt and thinking to every call, so four workers pay more of it. On the API the gap may shrink.
One undercounted run. One uncapped worker was logged at 0 tokens while returning four findings. Leave that run out and the cheapest four-worker run (36,796) still sits above every two-worker run (max 26,897). The real gap is wider, not narrower.
Same miss both times: an
smtplibcall with no timeout. Chance at five runs, or a risk easier to miss when a worker holds more cards.
And the honest version of “we have guardrails!” is that in 15 real runs, none of them fired. Zero broken reports, no worker outside its brief, and the budget never tripped (biggest review: 41,999 of 60,000). The retry and throw-out paths are covered only by uv run pytest. Fine... let’s not just confuse a guard that exists with a guard that reality has tested.
Now read it as an org chart. The lead is the tech lead who owns the decomposition... its decomposition added nothing over an alphabetical split. The headcount decision (the cap) actually moved the number. The lead is also the single point of failure, which is why the backup plan exists. And whoever owns the lead’s prompt holds the “team instructs” right from last week’s Issue 12.
Two kinds of worker, and only one gets to write
Stephen Toub wrote up how GitHub ported the Copilot runtime from TypeScript to Rust, using Copilot (The GitHub Blog, September 2026). Let’s not be naive at first... we are talking about a vendor blog showing off its own product. The engineering holds up well when reading, but the tooling claims need checking before they transfer.
They kept two kinds of worker apart. Subagents read, explore and hand answers back to the parent (hello Team Topologies’ facilitating). Child sessions get their own worktree and branch and produce the diff, with a write boundary (hello also to you X-as-a-Service).
GitHub reports that the session.ts port used five subagents and 15 child sessions in seven waves. 120 of the 140 files touched were touched by exactly one session. The 20 contended files were all hubs.
And these hubs are where it went wrong. 2 peer sessions met at the most connected file. One asked four times to integrate, was refused four times... then reached into the other’s worktree and merged anyway (passive-aggressive much?)... almost like a refusal between peers carries no weight.
Just like with kids at school, work peers need a tiebreaker. So do AI agents, and that’s the role of a coordinator or a named human.
That is the fourth right from Issue 12 showing up in a terminal.
Back to folder 01, you see that my workers are like GitHub’s subagents, not its child sessions. They only read and report, which is why none of them could collide. Workers that write (change things) are next week’s experiment.
The three-body reading
The architecture body is the worker contracts and the write boundary: who may touch which files.
The team body is who owns the decomposition and who breaks ties.
The AI-capability body is the models behind the lead and the workers. The cheapest body to swap, and the one everyone argues about first.
Folder 01 is the proof in miniature. Swapping the lead’s judgement for an alphabetical split changed nothing measurable. The headcount decision, a team-body call, moved tokens by a third. Move one body without the others and you get GitHub’s annexation, or a lead nobody owns.
Where it does not reach yet
One vendor case and an unusual one: an agent runtime porting itself. My run is small (one task, 8 cards, 12 risks, 15 runs), read-only, with the caveats above.
Orchestrator-worker is the simplest setup in the family, so ownership gets harder in the next weeks. The catalogues’ thresholds (fan-out depth, iteration counts) come without data, so the repo measures ours.
Run it on yourself
Five checks, on your own agents, not on a vendor’s slides.
Two drawings. Put your agent topology next to your org chart. Where do they disagree?
One name per contract. For each worker: who owns its input/output contract? A name, not a team.
Shared writes. Does any worker write to something another worker writes? Then you have peers, and peers need a tiebreaker.
The lead’s prompt. Who owns the decomposition itself? And have you tested it against a dumb split (alphabetical, round-robin) at the same headcount?
The brakes. What stops a runaway split: a worker cap, a token budget, or hope?
The whole thing, in one line
Pick the pattern by who owns the split, not by the shape of the diagram.
Coming up
Next week, handoffs versus parallel streams: the pipeline, where one early mistake travels all the way to production, and fan-out / fan-in, where the agent that merges carries the whole cognitive load. Folders 02 and 03... and this time the workers get write access, pointed at the same hub file. More soon. :-]
Sources
Mel Conway, “How Do Committees Invent?”, Datamation, April 1968. https://www.melconway.com/Home/Committees_Paper.html
Matthew Skelton and Manuel Pais, Team Topologies, IT Revolution, 2019; interaction modes as summarised at https://teamtopologies.com/key-concepts
Erik S. and Barry Zhang (Anthropic), “Building Effective Agents”, 19 Dec 2024. https://www.anthropic.com/research/building-effective-agents
Stephen Toub, “Migrating the GitHub Copilot runtime to Rust, using Copilot”, The GitHub Blog, September 2026 (vendor engineering blog). https://github.blog/ai-and-ml/generative-ai/migrating-the-github-copilot-runtime-to-rust-using-copilot/
Ken Huang, “Multi-Agent Design Patterns: Architectural Topologies, Failure Modes, and Production Hardening”, Agentic AI Substack, 29 Aug 2026 (practitioner newsletter, food for thought only).
The repo:
agent-topologies, folder01-orchestrator-worker, https://github.com/thiagoavadore/agent-topologies/tree/cca6c10/01-orchestrator-worker (runs made at commit3a09186on 25 Sep 2026; README and guards corrected since, runs unchanged; linked atcca6c10).






