Your bottleneck doesn't burn tokens
AI usage mandates put a jetpack on the one step of delivery that was already fast. Goldratt named the mistake in 1984, and the token dashboard is now making it at scale.
The 1:1 was going fine until the screen share.
It is a dashboard: AI usage by engineer, tokens per week, sorted descending. The manager scrolls to a name near the bottom of the list, and the question is almost kind. “Why is your number so low?”
Two desks over sits an engineer whose number is magnificent. He wears that dashboard the way other people wear a marathon medal. And in the same team’s repo, the oldest pull request has been waiting for a reviewer longer than either of them has been tracking tokens.
I have watched versions of this meeting more than once. The company changes, the tool changes… the dashboard survives.
Everyone in that room believes they are doing AI transformation. What they are actually doing is maximizing the input to the one step of delivery that was never the problem.
The mandate says: use the model more.
The delivery data says: your bottleneck doesn’t burn tokens.
Issue 1 named the trap: sequencing.
Issue 2 named the destination: AI-native, not AI-ready.
Issue 3 put three bodies on the table that have to move together.
Issue 4 climbed inside the AI capability and asked for one number: the acceptance rate.
This week, as promised: the team body. What the org chart has to give up for the model to stay in the loop.
And for the first time, there is a kit. Near the end you will find a five-question team-shape test and a prompt you can paste over your own delivery data, today, with the tools your mandate already bought.
If you run it, subscribe and tell me what your Herbie turned out to be. Herbie is explained below, and he matters more than your token count.
The mandate arrived with a leaderboard
In April 2025, Tobi Lütke told Shopify that “reflexive AI usage is now a baseline expectation” and wired AI competency into performance reviews and hiring. A few months later, Brian Armstrong told the Cheeky Pint podcast that Coinbase bought Cursor and Copilot, gave engineers a week to onboard, and fired the ones who did not. Coinbase now says roughly a third of its code is written by AI, aiming to have it at 50%…
Let me be fair to the memos before I take them apart. Refusing to seriously try these tools in 2026 is a career decision, and a bad one. If the mandate stopped at “try it, properly, this quarter”, I would co-sign it.
But the mandate never stops there, because an executive who mandates something needs to measure it… and usage is the number that is easy to get. Seats. Active users. Tokens per engineer per week. Somewhere down the org chart, that number becomes a sorted list… that list becomes a 1:1 conversation and even a way to evaluate “performance”.
All this knowing that measured people have always done: they optimize the metric. Pipe the model into everything, let it summarize what nobody reads, regenerate until the meter looks respectable. Call it tokenmaxxing. Burning tokens is the one KPI an engineer can max without shipping anything and some wear it as a badge of honor. A leaderboard where the prize for winning is a bigger invoice.
Which is exactly where the story has now arrived. The same companies that mandated usage are staring at their AWS Bedrock, Anthropic and OpenAI bills and asking where the value went.
They mandated an input… they got the input…
And now they act surprised because input is not the same as value created :-D
Issue 4 called usage the vanity metric: the treadmill is plugged in, nobody is running. This issue is about something worse. Even if everyone runs, maximizing this particular input cannot move delivery, even in principle. To see why, you need a book that is older than everyone building those dashboards’ careers.
Herbie doesn’t type
In 1984, Eliyahu Goldratt published “The Goal”, a novel about a factory manager (which is a sentence I promise gets better). For the academices out there, it is not peer-reviewed… it is barely literature. At the same time, it is also the clearest thing ever written about why your AI mandate is not working.
The famous scene is a boy scout hike. The troop keeps stretching apart because one kid, Herbie, is slow. The fast kids sprint ahead, then wait. The line’s speed to the destination is Herbie’s speed, no matter how impressive everyone else’s pace looks. The fix is not to yell at the fast kids to walk faster. It is to put Herbie at the front, take the extra weight out of his backpack, and subordinate everyone’s pace to his.
Out of that hike comes the Theory of Constraints and its two brutal accounting rules:
An hour lost at the bottleneck is an hour lost for the entire system, and an hour saved at a non-bottleneck is a mirage.
An AI usage mandate is a jetpack for every scout except Herbie.
Because in software delivery, typing code was almost never Herbie. Ask where a change actually spends its calendar life. Waiting for review. Waiting for the meeting where two teams agree on the interface. Sitting in the handover between the team that wrote it and the team that operates it. Or worst of all, being rebuilt, because what shipped was not what the customer needed and finding that out took a full cycle. The constraint lives in the seams between people. It always did.
Now make code creation dramatically faster, company-wide, by decree. What happens? The queue in front of Herbie gets longer. That is the entire story of enterprise AI adoption in one sentence, and Goldratt priced it forty years before the first token was billed.
The data agrees, awkwardly
Two pieces of evidence, one small and sharp, one large and blunt.
The sharp one: Joel Becker, Nate Rush, Elizabeth Barnes, and David Rein at METR ran a randomized controlled trial (July 2025) with 16 experienced open-source maintainers doing 246 real tasks in repos they knew deeply. Before starting, the developers forecast AI would make them 24% faster. Afterwards, having lived it, they estimated it had made them 20% faster. The stopwatch said they were 19% slower.
Sixteen developers, early-2025 tools, unusually mature codebases: carry the number with care. But carry the shape of it everywhere, because the gap between felt speed and measured speed is the fuel every usage mandate runs on. If the person doing the work misreads their own speedup by nearly 40 points, the executive reading a token dashboard three levels up has no chance.
The blunt one: the 2025 DORA report (Google’s research program, vendor-funded, not peer-reviewed, and honest about both) finds that AI adoption now correlates with higher delivery throughput and with higher instability at the same time. More change ships, more change fails and the pile-up happens exactly where Goldratt said it would: review, testing, QA. Their respondents say reviewing AI output is harder than writing code, that time saved writing is re-spent auditing, and 30% of developers report little or no trust in AI-generated code. DORA’s own framing is that AI is an amplifier: it does not remove your bottleneck, it delivers more traffic to it.
I'd be reluctant to follow these studies as they are 2025 dated and data-bounded to that era… while we all know how much has changed since Opus 4.6 and the whole Agentic AI boom in the last 8 months. Nevertheless, I see and hear these arguments and feelings in every customer…
Which brings me to the promise I made last week, about the acceptance rate that climbs for a quarter and then quietly slides back.
Last week's edition argued acceptance rate is the honest number, and it is. Honest about the model. It is not honest about the system, because when the constraint sits downstream of the model, the number decays in one of two ugly ways. Either reviewers drown and start waving things through, so acceptance climbs while your change-failure rate climbs with it. Or reviewers hold the line and become the queue, so acceptance stalls, your best senior turns into a full-time auditor, and adoption slides back to whatever the team could actually digest. Nothing about the model changed in either case. The team shape did the deciding. That is what a team-shape problem wearing a model costume looks like from inside.
What the org chart has to give up
Goldratt’s method, compressed: find the constraint, get the most out of it, and subordinate everything else to it. Subordinate is the step where org charts flinch, because subordination costs status. Here is what it costs, concretely.
Give up the usage league table. Replace tokens-per-engineer with flow at the constraint: time to first review, time in review, age of the oldest open PR. You cannot manage Herbie with a dashboard that only shows the fast kids. Measure the queue, not the faucet.
Give up output-per-person as the status currency. If review is your constraint, then review capacity is first-class engineering work, not a tax on the seniors’ “real” job. The engineer who reviews fast and well is not underperforming on code output. They are carrying Herbie’s backpack, and the career ladder has to say so out loud, or nobody sane will keep doing it.
Give up the single gate. One architect who must bless everything, one platform team every change routes through, one codeowner file that funnels a whole repo to one person: these are constraints you built on purpose and forgot about. Redraw ownership so one team can take a change, including the model’s share of it, from start to shipped without leaving the room. This is the argument of Skelton and Pais’s “Team Topologies”: teams are the unit of delivery and their cognitive load is a budget you design for, not a personal failing.
Give up the ceremony. Alignment that lives in standing meetings and handover documents is the constraint wearing a calendar costume. Every handover is a queue. Every weekly sync that exists because two teams share a seam is Conway’s law sending you an invoice. Remember from last week that the model reads your architecture like an honest new hire? Your meeting load is the same reading, of the team body this time.
Notice what is not on this list: a bigger model, a better prompt, one more token. The 8% to 80% story I have carried since the first Substack post was never a model story alone. That step only held at 80 because the team around it changed shape at the same time, ownership included. The model got it to possible. The org chart got it to Tuesday.
Run the team-shape test
Five questions, yes or no. Count your yeses.
Is your primary AI metric a usage number (seats, active users, tokens), viewable per person?
Take the last meaningful change your team shipped: did it spend more calendar time waiting (for review, for approval, for a meeting) than being built?
Can you name, without thinking, the one person or team most changes queue behind?
Has an AI metric (usage or acceptance) climbed and then slid back within two quarters, with no model change to blame?
If every engineer’s AI usage doubled tomorrow, would the first thing to break be a human queue: review, QA, or a release approval?
0 to 1 yes: your constraint may genuinely be code creation. Mandate away, you strange and fortunate shop.
2 to 3 yeses: mixed picture. Start measuring the queues this week, before the next dashboard review measures the wrong thing again.
4 to 5 yeses: you have a team-shape problem wearing a model costume. Your org chart owes Herbie an apology, and the next section is for you.
The prompt
Paste this into whichever assistant your mandate bought you, and feed it real data. It takes about twenty minutes to collect and one meeting to act on.
You are helping me find my software delivery constraint. Be blunt. Do not flatter me.
I will give you three inputs:
1. My team's last 20 merged PRs/MRs, each with four timestamps:
opened, first review comment, approved, merged.
[paste]
2. One typical week of the team calendar: each recurring meeting,
its length, and how many people attend.
[paste]
3. Our AI usage snapshot, if one exists: seats, active users,
tokens or suggestions per week. If none, I will say "none".
[paste]
Compute, per PR: time to first review, time in review, and time from
approval to merge, each as a share of total open-to-merge time.
Use medians, not averages.
Then answer three questions:
- Where does a change spend most of its calendar life: being written,
waiting for a human, or waiting for a meeting?
- Who or what is my Herbie: the single queue the median change waits
in longest?
- If my team's AI usage doubled tomorrow and nothing else changed,
which queue grows first?
End with: the one team-shape change most likely to cut the median
open-to-merge time, and the one metric I should stop reporting.
The model will happily run this analysis. It is, after all, already in the room.
A mandate can make everyone use the model. Only the org chart can make it matter. Feed the constraint, not the meter.
Coming up
Next week I open the field notebook: what it actually looks like when a shop moves its constraint, all three bodies at once, and the before-and-after numbers that survived contact with reality.
More soon. :-]
Sources
Because none of this comes only from my head. Every link checked.
The constraint, Herbie, and the mirage: Eliyahu M. Goldratt and Jeff Cox, “The Goal: A Process of Ongoing Improvement“ (North River Press, 1984). A business novel, not peer-reviewed, and still the sharpest operations book on this list.
The Shopify mandate: Tobi Lütke, “Reflexive AI usage is now a baseline expectation at Shopify“ (internal memo, self-published on X, April 2025). Primary source.
The Coinbase mandate: Brian Armstrong on Stripe’s Cheeky Pint podcast, as reported by CIO, “Fire any developer who doesn’t use AI” (September 2025). Press coverage of a podcast, not peer-reviewed.
The perception gap: Joel Becker, Nate Rush, Elizabeth Barnes, and David Rein, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity“ (METR, arXiv 2507.09089, July 2025). Randomized controlled trial, preprint, not yet peer-reviewed.
Throughput up, instability up: the 2025 DORA report and its “Balancing AI tensions” insight (Google Cloud’s DORA research program). Large-scale survey, vendor research, not peer-reviewed.
Teams as the unit of delivery: Matthew Skelton and Manuel Pais, “Team Topologies“ (IT Revolution Press, 2019). Practitioner framework, not peer-reviewed.
I make this Substack thanks to readers like you!





