Oct 7, 2026
Two shows of hands
On 23 September, David Heinemeier Hansson walked on stage in Austin to open Rails World, a conference for about 1,200 Ruby on Rails developers. Hansson, known to everyone in software as DHH, created Rails more than twenty years ago. For most of that time he has been one of the loudest defenders of programming as a craft.
That morning he told the room that his company, 37signals, had gone “pencils down”. Writing code by hand there, he said, “is now an exceptional state. It is like seeing a bug in Sentry.” Then he asked who in the audience still wrote a material amount of code by hand each week. About five hands went up. One attendee later put the number nearer 30. Either way, it was a handful of people in a room of more than a thousand.
The recording has been watched almost half a million times. The most-liked comment under it, with 1,900 likes, reads: “This is the most confusing funeral I’ve ever experienced.”
The same conference closed with a keynote from Aaron Patterson, who sits on the Ruby core team and works on the language’s compiler and runtime. He asked the audience a different question. How many of them enjoyed receiving AI-generated pull requests? Two or three hands went up.
Put the two polls side by side. Almost nobody in that room said they still wrote a material amount of code by hand. Almost nobody enjoys receiving AI-written pull requests to check. The machines can now produce change faster than organisations have learned to check it. That gap sits at the centre of the fight over AI coding, and managing it is a job for leaders.
The story everyone tells
On one side is a vanguard that says the old way of building software is finished. Boris Cherny, who created Claude Code at Anthropic, said in January that all of his code had been written by AI for more than two months: “I don’t even make small edits by hand.” Anthropic puts the share of AI-generated code across the company at 70 to 90%. Ryan Dahl, who created Node.js, posted the same month that “the era of humans writing code is over.”
The converts are just as telling. Andrej Karpathy, a co-founder of OpenAI, called AI coding tools “slop” in October 2025. Two months later he wrote that he had “never felt this much behind as a programmer.” By February he had a name for the new way of working, agentic engineering, where “you are not writing the code directly 99% of the time, you are orchestrating agents who do and acting as oversight.” DHH told Lex Fridman in the summer of 2025 that he did not let AI write code directly. By January he said his position had “flipped”.
This is no longer a small group. JetBrains surveyed more than 15,000 professional developers between May and July this year. About one in five now writes no code at all without AI help. About 22% get more than 80% of their code from agents, and senior developers are more likely to work this way than juniors.
On the other side is a loud chorus that has turned “vibe coded” into an insult. The most-liked comment under Fireship’s video about DHH’s keynote, which has been watched 1.8 million times, reads: “I write artisanal, hand-written, grass-fed, free-range code.” It has 7,300 likes. Open source projects have started closing AI-written contributions on sight. The Python type checker mypy, for example, now states that pull requests from new contributors “that are mostly generated by LLMs with little human input will be closed.”
The tidy reading of all this is that the second group has fallen behind. They have not learned the tools, so they dismiss what the tools produce. Their managers know even less, so they take the engineers’ word for it. The company slows down, and nobody notices until a competitor that made the switch pulls ahead.
That was my own reading when I started looking into this. I have made a version of the argument before. In “The Man With the Red Flag” I wrote about the man Victorian law required to walk 60 yards in front of every steam road engine, waving a flag to hold it to walking pace. An engineer who sneers at vibe-coded work looks a lot like that man.
The evidence does not support the tidy reading.
What the sceptics are actually saying
The place to look is where engineers argue in public, in the comment sections under the most-watched videos about AI and coding. The tone is mockery, grief and complaints about quality. Read closely, though, and many of the people sneering say they use the tools every day.
The clearest case is ThePrimeagen, a former Netflix engineer and one of the most-watched programmers on YouTube. For years he was one of the sharpest critics of AI coding. In September he posted a video called “Am I Alone”, and it has been watched more than 900,000 times. The line from it that viewers quote most is: “I have had to let part of myself die to be able to run this fast.” The top comment, with 3,100 likes, comes from someone in the same place: “Am I more productive? Sometimes. Am I burning out at a record pace? Yes.” These are heavy users describing what the tools cost them.
The surveys say the same thing at scale. In Stack Overflow’s 2025 survey of more than 49,000 developers, 84% were using or planning to use AI tools. Yet 46% said they distrusted the accuracy of what those tools produce, against 33% who trusted it. Experienced developers were the most cautious: 2.6% highly trusted AI output and 20% highly distrusted it. Google’s DORA research found the same split. Of its respondents, 90% used AI at work and 30% had little or no trust in the code it generated.
Adoption has gone further than the trust. METR, a research group that runs controlled studies of developer speed, had to change its study design in February because developers no longer wanted to work without AI. Between 30% and 50% said they were holding back tasks rather than do them by hand.
The critics and the users are often the same people, and what they complain about is specific. A post on Hacker News earlier this month asked whether anyone was producing good code with coding agents. Its author described the job as it now is: “I spend nearly the whole workday slogging through convoluted code riddled with footguns.” Under one popular video, a software engineer wrote that most engineers they talk to “view their jobs as simply janitors for all the AI slop.”
This group has also heard one argument so often that it now makes them angry. Lars Faye, a developer, posted a video in September called “people are lying about agentic coding” because he was “pretty tired of hearing ‘skills issue’ every time someone highlights the very serious issues that come with AI coding.” The most-liked reply says the failures of a technology meant to remove skill barriers keep being “explained by claiming you don’t have the skills to use it.” Anyone planning to tell engineers that their scepticism is a skill problem should know they have heard that argument many times, and rejected it.
The bottleneck moved
For most of the history of software, the scarce resource was the time it took a skilled person to write code. Testing, review and release were built around that pace. A team could only produce so many changes in a week, and its checks were sized to match.
AI has relaxed that constraint far faster than anything downstream of it. The JetBrains numbers and the room in Austin show how far it has gone. The tests, the reviews, the understanding of how a change fits the rest of the system and the judgement about whether it is safe to ship have not kept the same pace. The constraint has moved from writing code to checking it.
Seen this way, the contradictions in the debate line up. Engineers can use AI every day and still distrust its output, because they are the ones absorbing the checking. DHH and Patterson can both be right. A leader who measures how much code is being produced is measuring the part of the system that is no longer scarce.
In “The limiting factor in an AI-native company is people” I argued that the number of people who can supervise agents sets the pace of an AI-native company. The two shows of hands in Austin are that limit seen from the floor. Software is where this shift has become easiest to see. The same problem is likely to appear wherever AI makes production cheap but checking stays expensive, from contracts to credit memos.
What separates the teams that cope
Some organisations are getting real benefit from AI coding while others are breaking things faster. The difference lies in how they test and check their changes.
Google’s DORA programme has studied software delivery for more than a decade. Its 2025 report found that AI adoption now goes with higher throughput and better product performance. It also found that AI adoption still goes with lower delivery stability. Its explanation fits in one sentence: “AI doesn’t fix a team; it amplifies what’s already there.” Strong teams use AI to get better. Struggling teams find that AI makes their existing problems bigger. The mechanism is ordinary. AI raises the number of changes a team ships, and without good automated tests, version control and fast feedback, more changes means more breakage.
The companies that cope are redesigning the checking itself. At Anthropic, Claude reviews every pull request, and Cherny says “there’s still a layer of human review after it.” He still reads code himself: “I don’t think we’re kind of at the point yet where you can be totally hands-off.” Linear, the project management software company, wrote in September that “agents have made it exponentially faster to ship code, but validating those changes hasn’t quite kept up at the same rate.” It rebuilt its build and test pipeline. Its test coverage nearly quadrupled, and it still cut the time each pull request waits and roughly halved the machine time.
37signals learned the lesson the hard way. DHH told the Rails World audience about an early attempt to let designers vibe-code features in Basecamp 5. It left the architecture looking like “Swiss cheese”, and the company went back to writing code by hand for a while before trying again.
Karpathy chose his new term carefully. He called it agentic engineering, he said, to make clear that “there is an art & science and expertise to it.” In practice that expertise means writing clear specifications, building tests, reviewing changes and owning the whole system. Those are the things the sceptics in the comment sections say are missing. The engineer who complains about vibe-coded work is often describing the exact discipline that separates the teams DORA says are pulling ahead from the teams it says are falling apart.
Two ways to get it wrong
This is where leaders come in. The tidy reading assumed that managers side with their sceptical engineers and slow everything down. Where the evidence is on record, the push mostly comes from the top.
JPMorgan Chase now runs internal dashboards that label its technology staff as light, heavy or non-users of AI tools, and it has written AI proficiency into their goals. Business Insider reported that the new goals have been met with apprehension. Engineers describe the same pressure in the comment sections. The most-liked of those comments, with 909 likes, reads: “My boss tells me I’m an engineering manager but unlike a junior who can learn and get better I’m forever going to be micromanaging AI while losing my ability to understand the project.” The Pragmatic Engineer’s survey of about 900 engineers and engineering leaders, run in January and February, found that leaders were more positive about AI than engineers were.
Meta’s story is the clearest. It added AI token use to the metrics in its performance system, and people did what people do with a metric. They burned tokens to look productive, a habit that earned its own name, “tokenmaxxing”. On 2 September, WIRED reported, Meta took token use out of its reviews. “No one should use AI tools just for the sake of using them,” the memo said. By then 93% of Meta’s code changes were being made with AI help. The usage was already there. What the dashboard could not show was whether the work was any better.
In “You can’t delegate AI fluency” I argued that licences bought, courses completed and prompts sent measure activity, and that a leader who cannot tell activity from changed work will reward the wrong things. A token count on a performance dashboard is the clearest example of that mistake I have seen. It counts the part of the system that is no longer scarce and says nothing about the part that is.
Amazon shows what the same gap can look like from the inside. In March the Financial Times reported an internal briefing that linked a run of incidents, including a retail outage of almost six hours, to “Gen-AI assisted changes” and “novel genAI usage for which best practices and safeguards are not yet fully established.” According to the FT, Dave Treadwell, who runs Amazon’s engineering group, told staff that junior and mid-level engineers would need senior sign-off on AI-assisted changes. Amazon disputed the report. It said only one incident involved AI-assisted tooling, none involved AI-written code, and it had not introduced new approval rules. Whatever role AI played in the incidents, the briefing shows Amazon facing the same question: whether its engineering safeguards are keeping pace with new ways of producing code.
A leader who cannot tell how the work is being validated is exposed from both directions. A sceptical engineer says “this is vibe coded, throw it away”. Sometimes the engineer is right, and sometimes he is the man with the red flag. A dashboard says usage is up, and it cannot show the review queue, the tests nobody wrote or the senior people signing off changes they had no time to read. Without fluency, both decisions are guesses.
Aaron Patterson gave the technical reason the checking matters in his Austin keynote. A compiler, he said, makes programmers a promise: whatever it does, the behaviour you can observe will match the code you wrote. “Your AI has not made this promise to you. There is no As-If Rule for your English.” A fluent leader understands this the way a good finance director understands that a forecast can be wrong in ways the spreadsheet will never show.
Leaders do not need to judge every change themselves. They do need to know whether the organisation has credible evidence that its AI-assisted engineering is working. That means knowing what good looks like, which controls are in place and which questions to ask, and being able to tell real results from adoption theatre. Technical expertise can be delegated. Accountability for knowing whether the system produces the intended result cannot.
“The Man With the Red Flag” ended by asking whether anyone in the building was fluent enough to tell the man with the flag to step aside. Sometimes, though, the flag is there for a reason. A fluent leader can tell the difference.
What a fluent leader does
None of this needs a leader who can write code. It needs a leader who asks better questions and measures the right things.
When an engineer says “vibe coded”, ask what exactly is wrong. Missing tests, changes nobody read, no clear owner, or an architecture that no longer holds together. If the engineer can name the problem, they have handed you a quality standard. Apply it to all code, whoever or whatever wrote it. If they cannot name it yet, ask them to turn the objection into something you can observe: a test, a failure mode, an architectural consequence or an example. Experienced people often sense a problem before they can describe it.
When a team says AI writes most of its code, ask how it knows the code is right. Ask what checks run before a change ships, how much of the checking is automated, and what has happened to the change failure rate since the team switched.
Measure outcomes the way DORA does. Track change failure rate, time to restore service, review time and rework. Keep token counts and AI usage out of performance reviews. Meta has already run that experiment.
Redesign the checking, then fund it. If agents double the number of changes, adding human reviewers alone is unlikely to keep up. Invest in automated tests, machine review and fast pipelines, and keep human judgement for the changes where it matters most.
Use the tools yourself on real work. The test from “You can’t delegate AI fluency” still applies: at least one recurring part of your own work that AI has materially changed, and that you can describe in concrete terms.
Where the next reviewers come from
This shift leaves one question open, and nobody yet has the data to answer it. Senior judgement is becoming more valuable at the moment junior engineers do less of the work through which that judgement used to be built.
In a randomised study of 52 engineers learning a new Python library, Anthropic found that those using AI assistance scored 17% lower on a test of what they had just learned than those who coded by hand. Its researchers warned that productivity gains “may come at the cost of skills necessary to validate AI-written code.” One study cannot show what happens over a career. But every company that leans on senior engineers to approve AI-assisted work is drawing down a stock of judgement it may not be replacing.
If juniors do less of the work that built that judgement, where will the next generation of reviewers come from? It is the problem I described in “The Ladder Is Gone”, now showing up inside engineering teams, and it needs an answer before the loss of that learning becomes the next bottleneck.
Back in Austin
Both keynote speakers in Austin were right. DHH is right that for a growing share of developers, writing code by hand is becoming the exception. Patterson is right that somebody still has to stand behind what the machines write. That will not always mean a person reading every line. It will mean tests, machine review and human judgement, arranged so that the organisation has real evidence a change is safe before it ships.
Most engineers already understand both points, which is why they sound so conflicted in the comment sections. The people who most need to understand them are the ones signing the budgets, setting the targets and deciding whose warning to believe. That is where the fluency gap sits, and closing it is a job for the leader.
Sources
Rails World 2026 Opening Keynote – DHH (YouTube, Ruby on Rails channel)
DHH’s Rails World 2026 keynote: Pencils down, now what (Sublime Coding)
Aaron Patterson’s closing keynote at Rails World 2026 (Global Nerdy)
DHH declares end of hand-writing code at Rails World 2026 (DevOps.com)
DHH has gone completely off the rails... (Fireship, YouTube)
Top engineers at Anthropic, OpenAI say AI now writes 100% of their code (Fortune)
The man who built Claude Code was changed by it (notes on Lenny’s Podcast with Boris Cherny)
When AI writes almost all code, what happens to software engineering? (The Pragmatic Engineer; Karpathy and DHH quotes)
Vibe coding is passé (The New Stack; Karpathy on agentic engineering)
How much code do developers really let agents write? (JetBrains Developer Ecosystem Survey 2026)
AI policies in popular open source projects (arXiv; mypy policy)
Am I Alone (ThePrimeTime, YouTube)
2025 Stack Overflow Developer Survey (AI section, including experienced developers’ trust)
Stack Overflow’s 2025 Developer Survey reveals trust in AI at an all-time low (Stack Overflow press release)
Announcing the 2025 DORA Report (Google Cloud)
We are changing our developer productivity experiment design (METR)
Ask HN: Is anybody producing good code with coding agents? (Hacker News)
The Collapse of AI Software Engineering (The Infographics Show, YouTube; comment quoted)
people are lying about agentic coding (Lars Faye, YouTube)
AI Isn’t Replacing Software Engineers: The Truth Is Much Worse (RobertElderSoftware, YouTube; comment quoted)
Big companies are turning AI into a scoreboard (Business Insider; JPMorgan dashboards)
AI tooling for software engineers in 2026 (The Pragmatic Engineer)
Meta pushes its new AI agent on employees, but eases off on tokenmaxxing (WIRED)
Meta ends AI usage metrics in engineer performance reviews (BigGo Finance; memo wording)
Report on Amazon’s internal briefing on AI-assisted changes, 10 March 2026 (Financial Times)
Amazon is linking site hiccups to AI efforts (CIO; includes Amazon’s dispute)
AI coding has made CI a bottleneck, so we reworked ours (Linear)
How AI assistance impacts the formation of coding skills (Anthropic)
YouTube view and like counts are as of 3 October 2026.



