So what happens when you orchestrate AI agents? Does it use more tokens (costs more)? If so, does it produce a gain in the output that will drive better outcomes? Last week we started with this and talking about Agents Topologies, by only giving them read access.
So what happens when agents can write? This is what we have this week. Three agents, each in its own copy of the repo, all pointed at the same hub file: one shared config file, platform.yaml, that all eight services inherit their timeouts from. I split the services so that two workers would want the same setting. Built to collide, on purpose.
Oh yeah, using the state of the art available to you: Sonnet 5.5 does the work everywhere, and Opus 5.5 only sits in the merge seat in the second experiment.
uv run 03-fan-out-partial-reducer/experiment.py and off they went...
Voilá! 20 runs later, across the two versions where they had a way around the shared file, the shared timeout was contested... 0 times. Zero.
Nobody collided. And not because they were polite... because they never needed the shared file. One worker wrote an override (its own value instead of the shared default) in its own service.yaml. Another hard-coded the timeout straight into the code.
Everyone met their own need and walked around the shared file, and the platform ended up with about 3 different timeouts across 8 services (3.0 and 2.9 on average, depending on which route was open).
So I closed the override route. They took the code route. I closed that one too... and only then did they start fighting over the file.
But Thiago, isn’t that just agents being resourceful? Somehow, yeah... but look at what I was doing the whole time: my agents weren’t avoiding each other. My architecture was giving them a way out.
In Issue 13 the workers only read, which is why none of them could collide. GitHub’s peer sessions, on the other hand, met at the most connected file... one asked four times to integrate, was refused four times and merged anyway.
This week my workers write... and I have two promises to pay: the pipeline, where one early mistake travels all the way to production, and fan-out / fan-in (several agents in parallel, one merging), where the agent that merges carries the whole cognitive load.
Both experiments are in the public repo, one folder each, with the code, the raw runs and a plain-language README.
One thing I did differently this time, and it is the trust move of the week: I wrote my predictions down and froze them in PREDICTIONS.md, together with the checker and its tests against fake fixes, before the first run. Where I was wrong, the README says “miss”. You will see a few of those below :-D
If this lands, forward it to whoever owns the shared config your teams keep walking around.
A check at the end can’t see where the mistake was born
The first experiment (code and full results): four agents in a row over the same eight services: extract the facts, classify the risks, prioritise them and write the remediation plan. Each one only sees what the one before handed over. No lead, no orchestrator... a chain of handoffs.
Then I break stage 1 on purpose in half the runs. Either I move a fact to the wrong service or I drop it entirely... so 3 versions: a) no checks at all, b) one check at the end, c) a check after every handoff.
Erik S. and Barry Zhang’s “Building Effective Agents” (Anthropic, December 2024) calls this prompt chaining and suggests “programmatic checks (see “gate” in the diagram below) on any intermediate steps”... fine... this is that check, measured: 10 runs per version, 5 faulted and 5 clean, Claude Sonnet 5.5 in every role (checks included).
The check at the end caught 0 of 5 injected faults... none. By stage 4 the plan was perfectly consistent with its faulted input, so the last check looked at it and saw nothing wrong... by then the mistake was just the input.
The check after every handoff caught 4 of 5, every single one of them at stage 1, the stage that owns the mistake! It let one dropped fact through: an admin console open to the world (ouch!).
The price? In my runs, the every-handoff version halted 3 of 5 clean runs (2 wrongly, 1 rightly, a split decided by a script, not by a model) and used 46% more tokens than no checks at all (42,308 against 28,975 per run). The checks alone were 53% of its tokens. So yes, you pay... in tokens and in false alarms somebody has to go and look at.
Now the surprise! A fact moved to the wrong service healed itself downstream in 3 of 3 runs with no check at all. Huh? In the end, it still carried its file path and the later stages tied it back to the right file. A dropped fact had nothing left to tie to and reached production 2 of 2 times.
Facts that carry their provenance survive handoffs. Missing facts don’t announce themselves.
Now read it with the org chart in mind. Matthew Skelton and Manuel Pais’s Team Topologies (IT Revolution, 2019) says stream-aligned teams have “no hand-offs to other teams for any purpose”. A pipeline of agents is a chain of hand-offs, the exact thing I, and every Team Topologies practitioner, tell teams to avoid... but how can we not when we are talking about tokens explosions and missing context? So, when we must have hand-offs, we need to put the check where the work is born and accept the bill.
In Issue 12 I asked who owns each rule. Here the question is who owns a mistake. Every fault I injected was born in stage 1, the agent that extracts the facts. The check after every handoff caught it right there and blamed stage 1: the right owner. The check at the end never fired. And if it had, it could only have pointed at stage 4, the agent writing the plan, which did nothing wrong. It just inherited the mistake.
Nobody collided until the platform was the only way
Now the second experiment, where the cold open happened (code and full results). Three agents with write access, each fixing its own share of eight services.
All 8 read their defaults (timeouts, rate limits, backups, minimum instances) from one shared config file, platform.yaml. I split the work so it would collide as the two services that need a longer timeout went to different agents. Thus a script (not a model) scores the result by parsing and runnning the files... then it counts what got fixed and what a shared change broke.
But ofc, there were 2 traps inside:
pricing-engine needs a timeout of at most 300 ms and has no agent of its own.
payments-gateway needs at least 20 requests per second. A shared default that is right for the worker who set it can break a service nobody is looking after.
Claude Sonnet 5.5 as the workers, Claude Opus 5.5 in the merge seat, and each version ran 10 times. This is what we got:
Local overrides allowed. Shared timeout contested in 0 of 10 runs. 3.0 distinct timeouts across the 8 services, on average.
No overrides, but a timeout in code allowed. 0 of 10 again. 2.9 distinct timeouts.
The platform rule. Timeouts must come from the shared file, and a CI-style check reverts any file that hard-codes one. Contested in 6 of 10 runs, in every version with the rule.
Same models. Same prompts, apart from one neutral line per step stating the rule (no “please coordinate with your colleagues” nudges, I promise). And the collisions went from 0 to 6 of 10.
So it seems that contention is a property of the architecture, not of the agents. They collide only when the shared file is the only seam. Before that, they fork the platform quietly and you find out six months later when somebody counts the timeouts.
With some scientific honesty, the platform rule is my design choice. The dry runs never collided on the timeout without it, so I imposed it, and I’m saying so. The versions with the rule tell you what happens once the shared file is the only way... not how often that is true in your shop.
And this is Issue 12’s “platform enforces” right, as code. The rule is what made the shared default real. Before it, “the platform default” was a suggestion, nothing more than that...
The merger has to speak for the service that isn’t in the room
Once they collide, somebody has to merge. And that is promise 2 from last week, remember? The agent that merges carries the whole cognitive load... so let’s see what that load actually is. 4 ways to own the merge, all under the platform rule, 10 runs each (3 for me):
No owner, first change wins. pricing-engine broke in 10 of 10 runs. The shared timeout ended up above 300 ms every time, and nobody could grant pricing-engine an exception.
A supervisor (Opus) that resolves conflicts. broke in 4 of 10. And here is the part I did not predict: all 4 came from runs with no conflict. One worker raised the shared timeout alone, nobody disagreed, so the supervisor was never asked... and the uncontested value broke a service no worker owned. In the 6 runs that had a conflict? Zero breaks.
A platform owner that applies every change to the shared file itself. 0 breaks in all 10 runs. Its usual move: keep the shared default at 300 ms (9 of 10 runs) and grant the slow callers their own overrides (at least one in 6 of 10 runs, both in 5). The cost: the fewest problems fixed of the versions with the rule, 8.4 of the 11 planted ones against 8.6 and 8.7, because it also said no.
Me, in the merge seat. 3 runs. pricing-engine broke in 3 of 3, and I granted 0 overrides. I sat there with every service’s needs on the screen and picked between the two loud workers... how? I shared some text in the screen, some files to a free version of Copilot, asked to explain to me the problem and take a decision... what I know a lot of people do and say they are “using AI to take architectural decisions”.
So what is the cognitive load, Thiago? It’s not the conflicts. Those are the easy part... two values, two reasons, pick one. The load is the services nobody is speaking for. A conflict resolver only sees disagreements. The owner sees every change, including the quiet ones that break the service that isn’t in the room.
Team Topologies has a name for this as well and that is only “managing team cognitive load”. That owner is the platform team and its job is to carry that load so the stream-aligned teams don’t have to.
My answer key was wrong again, twice
Last week’s first lesson was “whoever writes the answer key decides the result”. This week I got to relearn it... twice...
First, the checkpoint was wrong the first time. In my first run, the check-after-every-handoff version rejected 5 of 5 clean runs. Five!
Great check, until you read the reasons... I never told the stage-2 check what stage 2 was allowed to drop, so it judged stage 2 against a bar the stage could never meet. Re-scored, 4 of those 5 rejections were false.
The fix was giving each check its own stage’s contract. Then I re-ran all 30 and kept the first run’s results in the repo, in plain sight.
Hmm... wait. Was it “whoever writes the answer key” or “whoever writes the checker”? It turns out that both... the check is part of the answer key.
Second, one planted risk (a problem I hid on purpose) was unfixable. The workers declined to give the maintenance scheduler a second instance in 5 of 5 trial runs, with the right reason: its data lives on the VM’s local disk, so two instances would split it. They were right and I was wrong. The risk came out of the answer key before the real runs (11 risks, not 12), written down in the checker’s notes.
And the frozen predictions have two misses worth owning. I predicted the end-only check would catch some faults (it caught none), and that the supervisor would have near-zero breaks (4 of 10, for the uncontested-change reason above).
Maybe I should stop writing answer keys :-P ... or maybe this is the job. A frozen prediction turns “I knew it” into “I was wrong here, and here is why”.
The three-body reading
The architecture body is the seams. Is the shared file the only route? Does each fact carry where it came from?
The team body is who owns the merge and who checks at the source.
The AI-capability body is the models. Unchanged across every version this week, and still the thing everyone argues about first.
The merge experiment is the proof. Same models, same instructions to the agents, and the service that broke went from broken in 10 of 10 runs to broken in 0 of 10. The only thing I changed was who owned the merge. The AI capability didn’t move at all.
Where it does not reach yet
One task, 10 runs per version, Claude models only, one backend. My test harness applies the agents’ edits, so these are not tool-using agents exploring a repo and committing on their own. The traps and the platform rule are mine, disclosed and pre-registered, which makes the bias visible... not gone. The human version is one person that built software before but inputs things into Copilot, talks a bit with it and pastes the content somewhere else...
Run it on yourself
Five checks, on your own agents. They work on your teams too, the checks don’t care.
The first check. In your agent pipeline, where is the first check? If it’s at the end, it can tell you something is wrong, not where.
Provenance. Does every fact or artifact carry where it came from (file, line, source)? Provenance is what let a misattributed fact heal itself.
Escape routes. Which shared config do your agents (or teams) write to, and what local routes exist around it? Count your distinct timeouts.
Every change. Who sees every change to the shared file, not just the conflicts?
The empty chair. Name a service with no owner in the room. Who speaks for it at merge time?
The whole thing, in one line
Check where the mistake is born and give the merge to someone who speaks for the service that isn’t in the room.
Coming up
Next week, one level up. Leads managing other leads, and agents that talk only through a shared message channel. Two questions: should the agent layers match your team layers, and who looks after the shared channel? It’s Conway’s law run backwards: design the team shape you want, and let the architecture follow. More soon. :-]
Sources
Erik S. and Barry Zhang (Anthropic), “Building Effective Agents”, 19 Dec 2024. https://www.anthropic.com/research/building-effective-agents
Matthew Skelton and Manuel Pais, Team Topologies, IT Revolution, 2019; stream-aligned teams and team cognitive load as summarised at https://teamtopologies.com/key-concepts
Stephen Toub, “Migrating the GitHub Copilot runtime to Rust, using Copilot”, The GitHub Blog, September 2026 (vendor engineering blog). https://github.blog/ai-and-ml/generative-ai/migrating-the-github-copilot-runtime-to-rust-using-copilot/
The repo:
agent-topologies, folders02-pipeline-checkpointand03-fan-out-partial-reducer, withPREDICTIONS.md(frozen 5 Oct 2026). https://github.com/thiagoavadore/agent-topologies/tree/7888d48 (results at commit7888d48, pushed 6 Oct 2026).





