In the loop. Not in the room
A model in your delivery pipeline is two things at once, a build lever and an architectural mirror. But what share of its work you ship without redoing it?
The slide says the AI capability is delivered. Seats are bought, the usage graph climbs, and someone has drawn a green checkmark next to the third body. On paper, the model is in your pipeline.
Then you look at where the work actually happens. The model drafts, and a person quietly rewrites most of it before anything ships. It suggests, and the suggestion gets read, sighed at, and closed. The tool is open on every screen and the job is being done the same way it was 6 months ago.
That model is present. Present is not the same as in the loop.
Issue 1 named the trap: sequencing.
Issue 2 named the destination: AI-native, not AI-ready.
Issue 3 put the AI capability on the table as one of three bodies that move together.
This week I climb inside that body and ask the only question that matters once the model is bought: is it in the loop, or just in the room?
A model is in the loop only where you would ship its work without redoing it. Everywhere else it is in the room: adopted, visible, paid for and not actually doing the job.
No kit this week. Just the argument and the standing ask: if you have put a model into a real pipeline, subscribe and tell me where I am wrong.
The model as a build lever
The first thing a model in the loop does is change what is worth building. Not how fast you type. The size of the thing one person can now finish.
Every backlog has a graveyard: the work that was permanently parked because the payback never cleared the bar. The manual review step no one would fund a rewrite for. The migration that needed six months of a team you were never going to get. A model in the loop moves the economics, not just the ergonomics, and some of that graveyard work suddenly clears the bar. The item that was “too expensive to bother with” becomes a Tuesday.
That is the lever. It is quieter than the demos suggest and more consequential. It does not make your existing plan faster. It changes which plan is worth having, which means it reaches straight back into the architecture body and the team body, the same coupling Issue 3 was about. Treat the model as a faster autocomplete and you will use a build lever to shave minutes off typing, which is roughly using a forklift to carry your lunch.
The model as an architectural mirror
The second thing a model in the loop does is show you your architecture, whether you asked or not.
A model can only act where your system lets it act. Point it at a clean seam with a real contract, and it moves: it reads the boundary, does the work, hands back something you can ship. Point it at a tangle where the batch job, the shared database, and the nightly cron all reach into each other, and it stalls, hedges, or invents. It cannot find the edge of the thing you are asking it to change, because there is no edge.
So where the model reliably fails, it is usually not failing. It is reflecting. A model that keeps getting the same file wrong is often telling you that file does three jobs. That is Conway’s law with a faster feedback loop: Melvin Conway said in 1968 that a system’s design mirrors the communication structure that built it, and now the mirror talks back in an afternoon instead of a reorg. The model reads your architecture the way an honest new hire reads it, out loud, on day one, without the tact. Most teams experience that as the AI being bad at their codebase. It is worth asking, at least once, whether the AI is being accurate about their codebase.
The wall chart of the future cannot tell you
There is a genuinely useful map of this territory, and I want to credit it before I argue with it. Amalfitano and ten co-authors, in a 2026 ACM TOSEM roadmap, sort GenAI in software engineering along two honest questions: what is being augmented (your process or your product), and how autonomous it is (does it wait for you, or act on its own).
That gives four clean forms, the Copilot, GenAIware, the Teammate and the Robot. As a shared vocabulary that beats the usual “AI for SE” hand-waving, and I will happily use it.
Then the same paper closes with ten predictions for 2030, and the precision is where it gets funny.
70% of boilerplate autonomously generated
“Prompt Architect” a certified job title
40% of job postings requiring it
20% of teams in regulated industries carrying a dedicated role.
These are not ranges or scenarios. They are two-significant-figure forecasts for a field that cannot reliably tell you what next quarter’s model will cost. And in the acknowledgments, the authors note the predictions were brainstormed with Claude Sonnet 4.5, off a literature review filtered by two other models. A roadmap for GenAI in software, whose boldest numbers were generated by GenAI.
The taxonomy is real work, so I say this without malice: the map of the future is cheap to generate and getting cheaper, and it will never tell you whether the model is in your loop this week.
The gap between present and in-the-loop shows up one layer down too. Moonshot just released Kimi K3 with open weights, and plenty of people read “open weights” as “we can run this ourselves now.”
The weights are about 1.5 terabytes and need multiple top-end GPU nodes to serve. Downloadable, yes. Runnable on your very expensive macbook or even a server under your table, no.
And a model can top the public leaderboard and still not be the one a team ships with, because fit to the actual work, the cost of a wrong answer, and where it sits in the pipeline decide more than the benchmark does.
Route the cheap work to the cheap model, keep the expensive judgment where it earns its keep, and stop confusing a model you can name with a model that is doing your job.
Present is not in the loop, at the model layer and the org layer both.
The one number
So how do you tell present from in-the-loop without a slide? One number: the acceptance rate of the model’s work. The share of what it produces that you ship without a human redoing it.
Not seats. Not tokens burned. Not “queries this month,” which measures that people opened the tool, not that the tool did the job. Usage is the vanity metric of AI adoption - it goes back to when we started doing BI dashboards and competed which teams had more.
Acceptance is the outcome. A usage dashboard tells you the treadmill is plugged in. Acceptance rate tells you whether anyone is running.
This is not just my framing. Albert Ziegler and colleagues, in “Productivity Assessment of Neural Code Completion” (MAPS 2022, expanded in Communications of the ACM in 2024), found that the acceptance rate of a model’s suggestions predicted developers’ own sense of productivity better than the other metrics they measured. The thing you keep is the thing that counts.
Which is why the one number I carry from the earlier issues is an acceptance number. That manual review step that went from roughly 8% machine-handled to 80% was never “the model was consulted 80% of the time.” It was 80% of cases closed by the model’s output and shipped, no human redoing the work. The 72 points between those two states is the model moving from in-the-room to in-the-loop on that one step. Had the number instead been “80% of reviewers now have the AI tool open,” we would have been measuring furniture and calling it transformation.
Run the in-the-loop test on your own pipeline
You can tell present from in-the-loop in about 3 questions.
Can you say your acceptance rate today, for one real step, as a number? If the only figures you have are seats and logins, you are measuring the room, not the loop.
Where the model stalls, is that a model limit or an architecture seam? Point it at your worst file and watch. If it keeps missing, the file is probably doing three jobs, and the model just told you which three.
What would you ship from the model without redoing it? If the honest answer is “nothing yet,” the model is present, and you have a build-lever and a mirror problem to fix before you have an AI capability.
Run it on your own pipeline before the next planning cycle, not on a vendor’s demo after.
A model is in the loop only where you ship its work without redoing it. Everything else is furniture with a login.
Coming up
Next week, the body nobody puts a number on: the team. What the org chart has to give up for the model to stay in the loop, and why an acceptance rate that climbs for a quarter and then quietly slides back is almost always a team-shape problem wearing a model costume.
More soon. :-]
Sources
Because none of this comes only from my head. Every link checked.
The taxonomy and the 2030 predictions: Domenico Amalfitano, Andreas Metzger, Marco Autili, Tommaso Fulcini, Tobias Hey, Jan Keim, Patrizio Pelliccione, Vincenzo Scotti, Anne Koziolek, Raffaela Mirandola, and Andreas Vogelsang, “A Research Roadmap for Augmenting Software Engineering Processes and Software Products with Generative AI“ (ACM Transactions on Software Engineering and Methodology, 2026). Peer-reviewed. The four forms are in the classification section; the ten predictions and the note that they were brainstormed with Claude Sonnet 4.5 are in the concluding section and acknowledgments.
Acceptance rate as the signal: Albert Ziegler, Eirini Kalliamvakou, and colleagues, “Productivity Assessment of Neural Code Completion“ (Proceedings of MAPS 2022, the 6th ACM SIGPLAN International Symposium on Machine Programming), expanded as “Measuring GitHub Copilot’s Impact on Productivity” (Communications of the ACM 67, no. 3, 2024, pp. 54-63). Peer-reviewed.
The org mirror: Melvin E. Conway, “How Do Committees Invent?“ (Datamation, 1968). The origin of Conway’s Law.
Kimi K3 open weights and hardware footprint: Moonshot AI’s Kimi K3 open-weights release (July 2026), reported by The Week and others; the roughly 1.5 TB weight size and multi-node serving requirement come from the release and deployment coverage. Vendor release and press, not peer-reviewed.
I make this Substack thanks to readers like you!





