Brett Chalupa (from the great YouTube channel Brett Codes), a software developer with 20+ years, uploaded “I’m done coding with AI” about 3 weeks ago. The title sounds catastrophic for some people, but it’s a really interesting and personal story that covers craftsmanship, mental health and someone who is truly passionate about his work. Up until this last Sunday (Aug/30th) it has almost 700K views, with a comment section full of supportive and relatable experiences.
Brett is not an AI skeptic who never tried or “doesn’t know what he is doing”. He went all in for ~1y after his company practically ordered them to use it, with the argument “everyone needs to use AI or you will be left in the dust”... and for months (even) Brett felt like they were right and AI-coding was the future.
I think there is a chance that Brett will be used in discussions for something he didn’t intend to. To be the evidence against AI for coding, or the joke for those convinced that “AI for coding” is the only solution... the 2nd one will try to discredit people like Brett saying they didn’t do proper context engineering, they don’t know how to properly do agentic AI or claim he was not using the latest model. No one is completely wrong or right. Brett just shared his (well-reasoned) experience and he is happy with the outcomes so far... why can’t we be happy with that and support Brett (I’m asking those who will go against him)?
I prefer to listen to Brett and learn. He is not a hero neither a luddite... he didn’t enjoy being left out of the coding experience, felt distant from his purpose and was a mere stamper of AI-generated code... so, indeed: what is there to like about that routine, about that job? But was the model the root problem here? Or the harness? Let’s read it the way we read the lab reports last week: as a postmortem. It feels like his job was moved from making to verifying... and nobody redesigned the verifying.
While Brett’s video has many other threads to pick up from (mental health, chatbot advice, environmental cost, ...), we will stick to the review process this week, ok? :-]
Issue 1 named the trap: sequencing.
Issue 2 named the destination: AI-native, not AI-ready.
Issue 3 put three bodies on the table that have to move together.
Issue 5 found Herbie: your bottleneck doesn’t burn tokens.
Issue 6 brought the peer-reviewed witness: more shipped, same typing.
Issue 7 taught one repo and moved a fleet.
Issue 8 found the sandbox was not holding, and told you to keep the gate.
No kit this week, just the argument and the standing ask:
The cost line nobody budgeted
So what, Thiago? One developer got tired, recorded a video and the internet clapped... is that it?
Nah... not really. Brett’s openness displayed a pattern I had actually investigated in a more academic way since January this year.
I started my thesis really curious about Team Topologies and the impact of GenAI in the Team Cognitive Load of Software Architecture Practitioners... So I ran a three-round Delphi with 8 to 10 software and enterprise architects at Antwerp Management School. A disclaimer is they were all Europe-based, AI early adopters... and yes, a master’s thesis is not peer-reviewed, here it is if you want to check my homework. I came out with two facts that I recognized from Brett’s video:
8 out of 8 participants in the first round verified everything. Not one of them accepted an output without review. “Everything is treated as hypothesis, never conclusion”, one of them claimed. And these are people who use the tools every day and like them!
when the panel rated what had changed the most in their work, the verification burden came out on top, with the strongest consensus of the entire study. Not “new skills”. Not “I produce more”. The checking.
And here is the irony the panel handed me without noticing (a perk of doing consensus-building rounds with Delphi): they rated “new skills emerging” low. For them, verifying is the old expertise... stretched. Nothing new to learn, just a lot more of it. Which is exactly why nobody hires for it, trains for it or budgets for it. It “is not new”, so it does not exist on any plan. It just shows up in someone’s evening.
“But Thiago, those are architects! Coding is different!”... is it, though? Guangrui Fan and colleagues put 60 developers through the same Python tasks with and without an AI assistant (CHI 2026). With AI: 22% faster, more correct... and stress and fatigue still went up... what explained the rise? A lot more verifying of the AI’s output, task after task. The help did not remove the burden, it just moved it.
Brett said “not possible and exhausting”. The panel said “verification burden”. The lab said “fatigue”. Different words, same thing :-D
Now... who is paying for it?
The gate, from the other side of the MR
Remember the six weeks from two issues ago? Teach the AI one repo, validate the lesson, run it across the fleet so that every change lands as a merge request that the repo owner reviews. The real experience I shared in a major European airline ended with 3 repos fully migrated, the rest ~70% underway, everyone happy. I told that story from the engine room... and last week I told you to keep that owner-reviewed gate, because that is where a wrong lesson stops before it spreads to the whole fleet, remember?
But now, let me tell the same six weeks from the reviewer’s chair...
By week 4 the MRs were arriving faster than anyone had planned for. Not one at a time, in batches. And the people reviewing them were the same one or two seniors who knew the codebase best, ofc... who else would you trust with it?
By week 6 I watched one of them open their 3rd MR of that afternoon. They approved it waaaay faster than the 2nd, because it looked like the 2nd. Same pattern, same lesson, green pipeline. Were they wrong? Probably not... that was the whole point of teaching the AI the pattern! But at that moment the gate stopped being a review and became a ritual (the same way we delete the PagerDuty alert that fired by mistake again...) and nobody had decided that. It just happened.
Did the gate hold? Yes, it held because two people held it, at a pace that was never in any plan.
I thought this was a “me” problem, a small-fleet problem, until someone counted it at scale. Hao He, Shyam Agarwal and colleagues at Carnegie Mellon and Stanford got access to the telemetry of a mid-sized software company that had publicly told its engineers to double their output with AI (“AI Writes Faster Than Humans Can Review”, a preprint from July, so not yet peer-reviewed, and one company only... but 802 developers and two years of pull requests, so not a survey either). And here is the thing: output did double. But are we looking for more output or more outcome? Let’s not go into that now, but don’t forget to read Mik Kersten’s Output to Outcome. Anyway, from He et al.’s paper there were 3 things I could not unsee:
per-reviewer load doubled, because the pool of reviewers grew a lot slower than the pool of MRs.
the share of MRs that any human reviewed fell from 89% to 68%.
and at the median reviewer, silent approvals roughly doubled while reviews with an actual comment stayed flat.
Their own sentence says it better than I would: “an enterprise AI mandate is a process-redesign problem, not a tooling deployment... it relocates work downstream rather than removing it.”
Relocates. Downstream. To the reviewer. My third-MR senior is in their data, and so is yours.
So, last week’s question again, people edition: you found the gate. Who is standing in it? How many MRs did they sign off last week? Did anyone count?
And if you are thinking about Herbie right now (the slowest hiker in the line, the one who sets the pace of the whole troop)... yes. Herbie is now the reviewer and your “agentic coding” just handed him 40 backpacks to carry.
Bainbridge already wrote this in 1983
From Brett’s video, what really scared me was his impression of engineers who “have basically deskilled themselves”... and I’ve seen that in some customers: when Claude, Codex or Cursor goes down, it’s almost like nobody knows how to develop anymore...
And Brett’s fear is not new. It even has an “official name” for the last 43 years.
In 1983 Lisanne Bainbridge published four pages in Automatica called “Ironies of Automation” and it became one of the most cited papers in human-factors engineering. Bear with me, as she was not writing about Software, but about control rooms and cockpits. Read her two ironies with your fleet in mind:
Automate the routine and you leave the human with the two jobs they are worst at: staring at an MR that is mostly fine, hunting for the one-line mistake... and then stepping in cold when it is actually there.
The skill needed to step in decays exactly because the routine that built it is gone. You do not get good at catching the wrong MR by approving right ones all afternoon.
So what did she suggest? Automate more and catch all mistakes is what a command-and-control manager usually aims for. WRONG ANSWER and not feasible!! Less automation then? Also, not good... her fix was more training and more deliberate practice for the person left watching. Brett is basically Bainbridge in a hoodie: the same thing she described forty years ago, in a paper about chemical plants.
So if that is what happens to the human at the gate... who is actually at the gate today?
Kacper Duma, Patryk Wróblewski and colleagues at Nicolaus Copernicus University looked at exactly that (accepted at EASE 2026, a preprint, and open-source GitHub repos only, so read it as a signal, not as your company): >33k merge requests written by coding agents, compared with human-written MRs in the same repositories, with a finding I keep repeating to people:
Of the human-written MRs that got any review, 25% were reviewed only by humans. What about agent-written MRs, how many were reviewed only by humans? 8%. Fascinating: same maintainers, same repos, and when the author is an agent the human almost never reviews alone. Either an agent reviews it too, or the human is busy steering the agent instead of judging the change.
So, who reviews the reviewers? In a growing share of repos: another agent. Or a human who is not reviewing the code anymore, they are steering the thing that wrote it (a quarter of the human comments on agent MRs are instructions to the agent, not opinions on the change). Which is a perfectly reasonable way to work... as long as someone, somewhere, still reads the diff (or a form of diff).
So here is my three-body reading, and it is not a very happy one. Every model release makes your agents more capable. That is the capability body, sprinting. None of those releases makes your reviewer’s week smaller... each one makes their job bigger and their practice smaller. That is the people body, hollowing out quietly. Last week the gap was permissions, egress, security in a specific blast radius. This week it is the same gap, wearing a person.
The table, the factory, and the inspector
Let’s try an image that has been stuck in my head since I watched Brett’s video.
A skilled woodworker makes tables in a way that shows he truly cares about the joints, the grain, the way the drawer slides... it takes him a week, and he is happy, because that is how he wants to make a table. Across the road there is a table factory. Same wood supplier, machines everywhere, a table every four minutes and they come out fine. Really, they are fine. Most of the tables in the world should come from that factory and this newsletter has now spent eight weeks helping you build one.
But look at the end of the line. There is a stool there, for the inspector. The person who knows what a good table is, who can tell a bad joint from three meters away... because they spent years making tables by hand. The factory needs that person more than it needs any single machine. And the factory is the one machine that quietly stops producing them, because the only way to learn what a good table is was making tables.
Brett walking off the floor is not the craftsman sulking about the machines. It is the inspector leaving. And the factory kept running that afternoon, the way factories do... nobody noticed, until some tables started to wobble.
But wait! Before someone sends this to their team with the wrong lesson attached: this is not a story about two kinds of developers. The ones who “really care” versus the ones who “just want the job done”. Brett was the factory guy for six months (”this is the future”, remember?) before the distance did its work on him. Same person, before and after. The axis is not who you are, and it is not about who or what is better. It is what you spend your week doing: making vs verifying what a machine made. Move anyone from the first to the second for long enough and they either walk or stop looking at the joints. The panel of my MSc said it in their own words when I asked them who the architect is, now that anyone with a prompt can produce an architecture: “everybody can generate an architecture now, which is cool. But the one taking responsibility of the decision is in the end the architect.”
Production got distributed but the Accountability stayed. A human still signs and is responsible.
So the question for your org is not “do we have craftsmen or factory workers?”. It is: which parts of your factory are (or should be) fully automated? And where is the inspector’s stool, who is sitting on it, and what did you do last quarter to make sure there is someone who can?
The operating-model answer, people edition
Those questions have answers... boring ones to be honest. Same as last week: not a tool to buy and no committee to form. 4 decisions to take and none of them is expensive, nor requires a tender to select the best “AI Frontier Model”...
1. Budget the review. Issue 5 already said that review capacity is real engineering work, not a tax on the seniors’ “real job”. Now put the number in the sprint. If your team can absorb (for example) 30 agent MRs a week with an actual human reading them, then 30 is the number, and the fleet runs at 30... not at whatever the agents feel like producing on a Tuesday night. We’ve all heard about WIP limits and the “make the work visible” part, so budget a team’s cognitive load the same way: you do not “hope” it fits.
2. Decide what gets reviewed how, and write it down. This one is my proposal, not something the panel voted on, so take it as such. A trust ladder per type of task. At the top, the stuff you spot-check: summaries, boilerplate, dependency graphs (the panel found those were sometimes better than the hand-drawn ones, by the way). In the middle, the stuff you read line by line: anything touching security, compliance, money. And at the bottom, the stuff you do not delegate at all: the calls that carry organizational judgement, politics, “who owns this”. Publish the ladder. Then “reviewed” means the same thing on every MR, instead of whatever the tired person at 17:45 decided it meant after they got pinged 10x that day with “can you review this pls?”.
3. Deliberate friction. Also my proposal, at least for engineering... but the mechanism has evidence behind it. AI-free design sessions. Attempt the problem yourself for 15 to 30 minutes before asking the agent. Treat the agent as a pair, not as a subcontractor you throw tickets at. Giuseppe Romeo and Daniela Conti went through 35 studies on automation bias (AI & Society, 2026, mostly medicine and HR, not software, so same caveat as before) and 3 things came out that your reviewers will recognize: a) people rubber-stamp a machine when they are overloaded, not when they are lazy; b) telling them “you are accountable for this” changes nothing; c) making them decide first, and only then see the machine’s answer, does. Your team will grumble about the friction and that means the (planned) friction is working. Pilots still fly by hand on purpose, and Bainbridge told us why.
4. Rotate the gate, and make “reviewed” mean something. If the same two people sign everything, the gate is a person, and people wear out. Remember the company that doubled its output: silent approvals doubled, commented reviews stayed flat. So: the log records who reviewed what, agent reviews included, and a silent approve-only click does not count as human review of agent code. A human who left a comment, or a human who actually ran it, owns the merge. Yes, this is last week’s “keep the transcripts” plank again, one level up.
So you think I’m getting boring? Because I keep saying the same thing, that “it’s not about the model or the subscription”? Well... too bad, as I’ll keep on that cadence. The “best model according to benchmarks” changes every few weeks. The reviewers are the people you keep... as long as you give them a job that is possible to do.
Run it on yourself
Before the next license renewal, before the next “AI adoption” townhall, and definitely before someone forwards you Brett’s video with “should we worry?” on top... three things you can check this week. On your own org, not on a vendor.
Count the signatures. How many agent MRs did your best reviewer approve last week? And how many of those got an actual comment from them? If nobody can answer that without writing a query first, the query is your first finding.
Find the last blank file. When did that same reviewer last write something from nothing, no agent, no autocomplete? If the answer is “a quarter ago”, congratulations, you have Bainbridge’s second irony on payroll... and you are paying a senior salary for a stamp.
Run the Tuesday test. If Claude, Codex and Cursor were all down next Tuesday, which of your teams could still ship on Wednesday? Name them. That list is your inspector count, and it is probably shorter than your org chart suggests.
Whichever one stings, you now know which body has been sitting still while the other one sprinted (and the architecture body watched, again). And you have a much better answer than “we trust our seniors” for when the board asks.
The whole thing, in one line
You kept the gate. Now go and check who is holding it... and whether they still can.
Coming up
Next week I go where this issue kept pointing and did not walk: the craft question. Brett is not alone, there is a whole “I’m done with AI” movement growing in the comment sections, and I think it is not about AI at all. What the people rejecting AI coding are actually telling you, and what to do about it before your best inspector walks off the floor. More soon. :-]
Sources
Directly cited:
Brett Codes, “I’m done coding with AI”, YouTube, 11 August 2026 (practitioner video; 686,525 views on 30 August 2026).
Brett Chalupa, “I’m done using AI”, Brett Codes blog, August 2026 (the written version of the video). https://brettcodes.com/im-done-using-ai/
Thiago A. de Faria, “From GenAI Practice Vacuum to Validated Guidance: A Delphi Study of AI-Augmented Reference Architecture Development”, MSc thesis, Antwerp Management School, 2026 (the author’s own three-round Delphi study with 8 to 10 practitioners, January to March 2026; not peer-reviewed). https://doi.org/10.5281/zenodo.18367977
Guangrui Fan, Dandan Liu, Lihu Pan, Rui Zhang, “When Help Hurts: Verification Load and Fatigue with AI Coding Assistants”, CHI 2026. https://doi.org/10.1145/3772318.3791176
Hao He, Shyam Agarwal, Yegor Denisov-Blanch, Pavel Azaletskiy, Sanmi Koyejo, Bogdan Vasilescu, “AI Writes Faster Than Humans Can Review: A Longitudinal Study of an Enterprise ‘2×’ Mandate”, arXiv:2607.01904, July 2026 (preprint, not yet peer-reviewed; one firm, 802 developers, 196,212 pull requests). https://arxiv.org/abs/2607.01904
Lisanne Bainbridge, “Ironies of Automation”, Automatica 19(6), 1983, pp. 775-779. https://doi.org/10.1016/0005-1098(83)90046-8
Kacper Duma, Patryk Wróblewski, Jagoda Bobińska, Julia Winiarska, Piotr Przymus, “These Aren’t the Reviews You’re Looking For: How Humans Review AI-Generated Pull Requests”, accepted at EASE 2026, short paper (arXiv preprint; GitHub open-source repositories only). https://arxiv.org/abs/2605.02273
Giuseppe Romeo, Daniela Conti, “Exploring automation bias in human-AI collaboration: a review and implications for explainable AI”, AI & Society 41, 2026, pp. 259-278 (peer-reviewed PRISMA systematic review of 35 studies; domains mostly outside software). https://doi.org/10.1007/s00146-025-02422-7
Further reading (not cited in the body, same argument from other angles):
Hao-Ping (Hank) Lee, Advait Sarkar, Lev Tankelevitch, Ian Drosos, Sean Rintel, Richard Banks, Nicholas Wilson, “The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers”, CHI 2025. https://doi.org/10.1145/3706598.3713778
Raja Parasuraman, Dietrich H. Manzey, “Complacency and Bias in Human Use of Automation: An Attentional Integration”, Human Factors 52(3), 2010, pp. 381-410. https://doi.org/10.1177/0018720810376055
Rafael Tomaz, Paloma Guenes, Allysson Allex Araújo, Maria Teresa Baldassarre, Marcos Kalinowski, “Impacts of Generative AI on Agile Teams’ Productivity: A Multi-Case Longitudinal Study”, FORGE 2026 (the Issue 6 study). https://doi.org/10.1145/3793655.3793728
Mik Kersten, “Output to Outcome”, IT Revolution, 2026 (practitioner book; Issue 2 callback). https://itrevolution.com/product/output-to-outcome/
Matthew Skelton, Manuel Pais, “Team Topologies: Organizing Business and Technology Teams for Fast Flow”, IT Revolution, 2019 (practitioner book). https://teamtopologies.com/book





