Slow Down to Speed Up: Gergely Orosz on What AI Is Actually Doing to Software Engineering
The Pragmatic EngineerIn this keynote at Craft Conference 2026 in Budapest, Gergely Orosz, author of The Pragmatic Engineer, asked whether AI coding agents are really making teams and companies more productive. Orosz relies on sources inside many large tech companies, and argued that coding has changed dramatically in the past six months, but that the results so far include collapsing quality, runaway costs, burnout, and in one case a self-inflicted crisis at Meta. Their advice, captured in the title, is to slow down: ship only what you can verify, keep learning, and put AI to work at the level of the system, not just the individual.
The Instagram "Exploit" That Wasn't Really an Exploit
Orosz opened with events from the week of the talk, which they called the worst in Meta's history. On Monday, Instagram suffered what they called "the goofiest ever" security breach. It had two steps. First, an attacker used a VPN to fake a location matching the victim, for example the US if the target was Barack Obama's account. Second, the attacker told Meta AI that their account had been hacked and asked it to send a verification code to an email address they controlled. Meta AI sent the code, and the attacker could take over the account. Orosz described it as effectively the first "zero-auth password reset," and said that to engineers it is simply a bug.
What puzzled Orosz was how this could happen at a company with such strong engineering culture. They listed what Meta has: what they consider the world's best automated rollout and canary system, many layers of verification, manual code review, and an Instagram trust and safety organization of close to 100 engineers. None of it stopped the bug. On Tuesday, Meta's chief information security officer emailed staff to say they were quitting, in the middle of an incident investigation (a "SEV" in Meta terminology) that had not yet concluded.
Orosz said they had talked to people on Instagram's trust and safety team, and that the details had not yet been reported publicly. According to those sources, the faulty code was written by AI and reviewed by AI, not by humans. Orosz argued that AI alone did not explain it. They named three contributing factors: token maxing, layoffs, and what they called, with the caveat that psychosis is a serious condition, "AI psychosis" coming from Meta's leadership.
Token Maxing, Layoffs, and Forced Reassignment
Orosz had written in April about "token maxing," a trend at companies including Meta, Amazon, and Uber, where engineers began to be measured on AI token usage and responded by inflating it. Meta had an internal leaderboard with status tiers such as "session immortal" and "token legend." Meta shut the leaderboard down in April, but according to Orosz AI usage still figured unofficially in performance evaluation. Engineers understood that a low token count was a bad signal, so they used AI for everything: rather than writing something by hand or reading documentation themselves, they asked the AI, partly just to burn tokens. AI is free for employees inside Meta, and people wanted higher bonuses. The code behind the incident was AI-generated, and AI reviewed it, in some cases "triple" reviewed it.
The layoffs made this worse. Meta laid off 10% of staff, about 8,000 people, on May 20, but the layoffs had been reported about a month in advance. Orosz said that during that month, worried employees increased their AI usage so their token numbers would not drop and mark them as layoff candidates. Instead of focusing on their work, including trust and safety work, people were focused on inflating their numbers.
The "psychosis" part concerned reassignment. Orosz said Instagram's trust and safety organization, built over seven or eight years and mostly based in London, lost about 40% of its people before May 20. They were told on a Thursday that starting Monday they would move to Alexandr Wang's AI organization to do manual AI data labeling: reviewing GitHub pull requests, adding tests, writing feedback, then moving to the next task. They were not given a choice. Orosz stressed that Meta had always treated engineers "like royalty" and let them choose their teams. They reported that close to 5,000 Meta developers now do manual labeling, and that employees joke this operation is bigger than OpenAI. After the layoffs and reassignments, most teams are less than half their former size, and some services no longer have anyone on call, which Orosz said had never happened at Meta before.
Orosz called this fully self-inflicted, coming from Mark Zuckerberg and Alexandr Wang, and said the message was that building a frontier model mattered so much that Meta would risk its business and its security. They described morale as worse than during the 2022–2023 layoffs, and added that in the US Meta records employees' screens and keystrokes to train AI. Engineers told Orosz they feel like discarded tools. Even people given large retention bonuses are interviewing elsewhere, not out of fear of being laid off but because they don't know whether they will be reassigned to data labeling, work they did not sign up for as professionals. Orosz said this might be the end of the Facebook engineering culture they had come to admire, built over 22 years. They acknowledged that Meta is an extreme case, but said it is happening right now.
Why "Everything Changed in the Last Six Months"
After the Meta story, Orosz laid out the plan for the rest of the talk: what has changed, a tour of what tech companies are doing, industry-wide trends, and advice for engineers and leaders.
To support the claim that coding changed dramatically in the past six months, they cited well-known skeptics. David Heinemeier Hansson (DHH), creator of Ruby on Rails, told Lex Fridman in October 2025 that AI wrote none of his code, partly because the models were not good enough. By February, when he appeared on Orosz's podcast, that had flipped: most of his code is now AI-written. Orosz emphasized that DHH cares deeply about software craftsmanship, is not paid by any lab, and told them that the models now write better code than he does. Simon Willison, whom Orosz described as one of the most quoted people on Hacker News and an independent engineer who experiments with every model, said that models released in November 2025, which Orosz named as Opus 4.6 and GPT 5.4, made agents genuinely useful. Given six months to absorb that shift, Orosz said, it is unsurprising that companies are now spending heavily.
Orosz then shared data from partners. Linear compared teams using AI agents to teams that don't and found that teams with agents ship five times as much code. Orosz noted that a similar increase previously took perhaps 20 years and this one took under two, while promising to come back to quality. Cursor reported that lines of code per developer per year rose about 2.5x, from roughly 3,000–4,000 to over 8,000, and that PR size tripled. Orosz multiplied these into roughly six times as much code, and acknowledged the audience's smiles: that probably means six times as many bugs. Cursor also reported a sharp rise in developers accepting AI changes without manual review, especially from January, when Orosz said Opus 4.7 and GPT 5.5 came out and companies started trusting Claude Code, Cursor, and Codex. Meta merging AI-reviewed code to production, they said, belongs to that same shift.
Anthropic and OpenAI: The Most "AI-Pilled" Companies
At Anthropic, Orosz spoke with Boris Cherny, creator of Claude Code, in February. According to Orosz, Cherny runs five parallel agents on a laptop at all times and ships 20 to 30 pull requests a day while leading Claude Code hands-on. Cherny told Orosz that PRDs, written planning documents, are dead at Anthropic and have been replaced by prototypes. Claude Code is 100% generated by Claude Code, and Anthropic overall is about 70–90% AI-written, with no target. People simply use it. Anthropic built Claude Cowork in ten days and it became a major commercial success. Orosz said sources at Microsoft told them Microsoft tried to build its own equivalent, and that Satya Nadella gave the team a month, but two and a half months later there is still nothing. Orosz used this to show that AI accelerates some companies while others, like Microsoft, stay held back despite having it.
At OpenAI, based on a conversation with the Codex team at the Pragmatic Summit in San Francisco, Orosz described an internal ChatGPT build with a "fix it" button: take a screenshot, say "fix this bug," and Codex generates a pull request that an engineer, or with safety nets even a non-engineer, can merge. AI code review happens in multiple tiers. Some code can reach production with AI review only, while critical-path code requires human engineers. Most developers run several agents at once. Orosz relayed the joke that engineers walk around with laptops slightly open so local agents keep running, and said that when they asked an OpenAI interviewee whether they had an agent running during the interview, the answer was "I had five." People attend meetings while their agents work.
Orosz noted that these are the most AI-committed people in the industry, who talk about AGI as a matter of when, not if. Most OpenAI staff barely write code by hand anymore. That changed around October. Some developers still wrote about 30% by hand then, and Orosz thinks that share is fading. The Codex team writes everything with Codex and told Orosz that taste, knowing what to build, is becoming very important. Codex also works on itself: because the team is mostly in one time zone, they kick off overnight runs where Codex tests itself and proposes improvements, which engineers accept or reject in the morning. In meetings and debugging sessions, they send voice notes to Codex at the start and have results by the middle or end. Orosz said it sounds like science fiction.
Cursor, Google, and Meta's Internal Tooling
Orosz visited Cursor's San Francisco office in October. They said Cursor went all in on agents in January, moving on from the tab-and-editor experience while still keeping it available. Cursor built its own coding model, Composer, which Orosz described as one of the few good coding models outside OpenAI and Anthropic, and notably cheap, which matters for costs discussed later. Orosz noted they have no affiliation with Cursor. Cursor runs tens of thousands of NVIDIA GPUs leased from Azure, AWS, and others. Inference used to be its biggest cost, but it now also trains models, turning it into a mini AI lab. Orosz mentioned that SpaceX seemed about to acquire Cursor. Everyone at Cursor is technical: Lee Robinson, who works in developer relations rather than as a software engineer, used Cursor to migrate all of Cursor's sites to a different CMS and published the cost.
At Google, everything is custom. The internal IDE, Cider, was web-based and is now a VS Code fork. Jetski is the internal version of Antigravity, integrated with the Piper monorepo. Critique handles code review, Code Search is best-in-class (Orosz said Sourcegraph was inspired by it), Borg is Google's Kubernetes, and Monarch its Datadog. Gemini is integrated throughout, so the internal experience is good. The problem, according to Orosz, is that Gemini is not as good as Opus or GPT 5.5. Engineers use Claude Code where they are allowed to, but only in limited contexts, so Google's AI adoption lags other companies. Orosz said the CEO has acknowledged this and Google is working on a better model.
At Meta, everything revolves around building a state-of-the-art model. Meta has an internal coding tool, Metamate, and a feature called "trajectories": GitHub-style commits show the exact prompts used. It was rolled out in December without warning people that prompts would be visible, so everyone could see, for example, staff engineers asking the AI to write a for loop. One developer told Orosz they began prompting in Polish so fewer colleagues could read it. Orosz believes Zuckerberg wants a model better than Opus 4.8, and that either Meta gets it within a few months or much of the company ends up deeply demotivated.
Uber: What Serious In-House AI Tooling Looks Like
Orosz, a former Uber engineer, used Uber to show how much in-house AI tooling a large company builds. Uber has about 3,000 engineers and about 20,000 other employees. A developer experience team of around 20 people, now effectively an AI experience team, has built:
- An internal MCP gateway for discovering and registering MCP servers.
- An Uber Agent Builder and Agent Studio, a no-code, drag-and-drop way for the 20,000 non-engineers to build agents, plus a registry for them. Orosz compared it to OpenAI's public Agent Builder.
- An AI CLI, which Orosz called "the Claude Code for Uber," integrated with internal systems and multiple models.
- Uber Minion, background agents at scale, similar to Cursor's background agents but integrated with Uber's monorepo and experimentation system. Orosz said developers who can use Claude Code often prefer Minion because it works better and faster. For example, it analyzes a prompt and suggests how to rewrite it for faster, cheaper, better results, which Claude Code doesn't do yet.
- A code inbox that surfaces which of the flood of mostly AI-reviewed changes actually need a person's attention, with smart assignment and SLAs that escalate to someone else if a reviewer doesn't respond within about a day, like on-call tooling.
- Risk profiles that flag changes that look risky.
- uReview, an internal AI code reviewer comparable to CodeRabbit or Sonar.
Orosz said other large companies are doing the same, each with a dedicated infrastructure org: Stripe (Minions, Toolshed, Blueprints, devboxes), Ramp (Inspect, Glass, Dojo, Sensei), Shopify (Sidekick, an LLM proxy, a dev MCP server), Airbnb, and others. If you're proud of wiring an agent into Slack, they joked, that's cool, but this is the next level.
Startups and Traditional Companies
Among startups, Orosz sees the familiar pattern: agents write and review code, and there's plenty of Slack-based creativity. At one startup that had raised a $70 million Series B, someone jokingly told an agent in Slack to "fix all bugs in the codebase." It came back having found four critical authentication issues that left a back door wide open. Orosz also sees startup engineers vibe-coding their own replacements for SaaS tools, which they read as engineers having fun rather than a real business trend, and which they don't see at large companies.
The surprise, for Orosz, was traditional companies. They lack Uber-style platform teams but are not really lagging. Cisco rolled out Codex to 18,000 engineers in January, when Codex was still small, and uses it for complex migrations. JPMorgan Chase built a multi-agent framework, specialized agents labeling customer interaction data with evals and judge-based aggregation.
Trend: Individual Speed-Ups Don't Add Up to Team Results
The first industry-wide trend came from Laura Tacho, formerly CTO at DX and now leading developer experience at AWS, whom Orosz messaged the night before the talk. Tacho told them that many organizations see individuals doing great but no improvement at the team level, because they treat AI as a personal productivity tool, which she calls "individual speed-up juice": email summaries, Slack automations, even code generation. The companies seeing team-level results start from a business outcome instead, such as deploying faster, shipping more features at the same quality, or improving quality.
Orosz offered Spotify as an example, based on lunch with its CTO about six weeks earlier. Spotify's bar for AI use is that quality must stay the same. As a result, it isn't seeing huge output increases, but it has built internal quality-checking tools and deliberately rolls AI out more slowly, a contrast with Meta. Tacho's framing is to build agentic systems that reduce handoffs, make information easier to find, and remove friction while maintaining quality, which Orosz said few companies do. Tacho maps companies on two axes: AI usage (individual vs. team) and decision-making (simple automation vs. agentic systems). Most companies sit at individual/simple automation, and most want to reach team/agentic systems. Getting there, Orosz argued, requires what Uber did: building integrated systems with your own engineers over time and at great cost. You can't simply buy Claude Code or Cursor and get it, whatever vendors say.
Trends: Token Maxing, Tool Addiction, and Vanishing Middle Management
Orosz said token maxing, burning tokens on purpose without shipping value, is going out of style but persists under pressure to look productive, especially at US tech companies that "don't really care about budget until they do." They said it has happened at Meta, Amazon, and Microsoft, which they said still runs an internal leaderboard. They also described the pricing of AI tools as somewhat addictive: start at a $10 or $20 plan, hit a generous limit, upgrade to $100 or $200, then feel pressure to use up the allowance, and eventually end up on API pricing. For some people, especially in the first months, prompting feels like gambling ("one more prompt"), with people sleeping poorly and waking up thinking about their agents.
Another trend is the flattening of middle management: managers laid off, reassigned to individual contributor work (as at Meta), or told to be hands-on and manage less, often with AI given as the reason. Orosz pushed back on the popular disdain for middle managers and directors. In their experience, good middle managers are very technical and could be hands-on but choose not to. They watch what's happening and intervene, for example noticing a run of outages and pulling people into a task force to fix the underlying system, where engineers alone might just pile on. Good engineering management improves engineering culture, and Orosz said that removing it will make culture decline, which they called "a fact as far as I'm concerned."
At the same time, CEOs and CTOs are returning to coding. Guillermo Rauch, founder and CEO of Vercel, told Orosz he sees executives coding again "with a fury," including public-company CEOs messaging him excitedly about using Vercel or Claude Code. Orosz pointed out the combination: fewer middle managers to protect engineers, and executives vibe-coding things they think are complete when they aren't.
Trend: The AI Bill Arrives
Orosz called rising AI costs a "mega trend" they noticed only a week or two before the talk and had written about for subscribers the previous week. On the morning of the talk, Sam Altman posted that AI budgets had become a big issue for some companies. Orosz said they asked contacts at OpenAI whether Altman reads their newsletter, and heard that someone had posted it in Slack and he had read it. They also mentioned a joke circulating on Reddit about a partner discovering $15,000 gone from a shared account.
According to Orosz, Anthropic has moved enterprise customers onto API pricing, so anyone who isn't a startup or individual no longer gets discounts. GitHub Copilot made a similar change on June 1, two days earlier, and users are furious that budgets, $200 or whatever they were, that used to last a month are gone in three days. Uber's CTO said in March that the company had burned through its whole annual AI budget. Uber now caps AI spending at $1,500 per engineer per month, after which engineers use free models. Orosz said their research shows many companies doing something similar, some capping as low as $200 before falling back to zero-cost models on Copilot. Costs that approach an engineer's salary are something no one wants to pay, "no matter what the AI labs say."
Craft Trend: Quality Is Dropping Everywhere
Turning to the craft itself, Orosz said quality has dropped across the board, starting with their own experience. For about a month, as a paying user, every time they opened claude.ai and started typing, a React lifecycle refresh wiped out what they'd typed. They screen-recorded it and tweeted about it. A product manager replied that it was great feedback and would be fixed. Orosz read that as an admission that Anthropic had no idea millions of people hit the bug daily and wasn't dogfooding its own product. They said a bank would do better, even though Anthropic did eventually fix it.
OpenAI bragged about building its Agent Builder in six weeks with one engineer using Codex, but Orosz said users hit many problems at launch and its forum filled with unresolved complaints. Three months later, one user wrote that they had been bullish but P0-type bugs weren't getting fixed and it felt like abandonware. AI made it faster to build, Orosz said, but not better. At AWS, an engineer let the internal Kiro coding tool make changes, and the agent chose to delete and recreate an environment, causing a major outage. AI-generated code also took down part of Amazon's flagship retail site. Orosz framed both cases as over-reliance on AI or not caring about quality, and said Amazon now requires a senior engineer to review any AI-generated change, because junior engineers would just approve it.
Orosz highlighted Dax Raad, founder of OpenCode, an open-source AI harness with nearly a million daily active users that has grown about 10x in four or five months. On Orosz's podcast the week before, Dax said the team was shipping more hacks where they should have redesigned systems from scratch, and that "our judgment is off." Dax also observed that no competitor is beating OpenCode by using AI better, admitted "I don't think we're using AI that well," and said the team is telling itself to use less AI, to think more, and to build fewer things that matter. Orosz noted the irony of an AI tool founder saying this, and suggested OpenCode remains one of the highest-quality harnesses precisely because it slows down.
Craft Trend: Everything Is Broken, and Reviewers Are Burning Out
Orosz's example of things breaking was GitHub, where two weeks earlier pull requests disappeared for 8 to 12 hours. An unofficial uptime tracker that aggregates reported outages estimates that GitHub doesn't reach even one nine, meaning some part of it is down around 10% of the time, which Orosz called "absolutely unserious." GitHub's leadership shared unpublished numbers with Orosz, attributing the problems to a 3x load increase over two years. Orosz said they don't buy it. They did not claim AI-generated code was to blame, but argued that if GitHub can't handle 3x growth in two years, with more to come, something is wrong, since startups absorb that kind of growth easily. Mario Zechner, creator of pi, told Orosz that software has become a brittle mess, with 98% uptime feeling normal and UIs full of strange bugs. Zechner conceded this predates agents but said it feels like it's accelerating. Orosz added that they recently had a serious software problem with Magyar Telekom, though they doubted AI was involved there.
The last craft trend is that AI "slop" is burying the engineers who still care. There are far more pull requests, mostly AI-generated, and most developers see that the AI reviewer passed a change and give it a thumbs-up without reading it. The minority who still review properly, catching bugs, pushing back, and spotting duplicated code, are overwhelmed and burning out. At performance review time they aren't rewarded, because they aren't seen as the people shipping features. Some quit. Dax told Orosz that OpenCode is hiring many of these people, who left because they were the last ones keeping things alive and no one cared. With engineering management gone or less hands-on, there's no one left to care.
Kent Beck summarized it for Orosz: "We're accumulating code faster than we accumulate trust." Code has to be understood and trusted, and there's no time for that right now. Beck also said AI amplifies experience, so seniors gain the most and judgment is rewarded. Orosz cited Hillel Wayne: the only people successfully generating working TLA+ specifications with AI are TLA+ experts who spelled out exactly what to generate. The same applies more broadly, Orosz said. A junior engineer who has never built an iOS app can prompt one into existence, but it won't be maintainable. Old patterns are returning as a result. Dax told Orosz that OpenCode relies on domain-driven design and verbose guardrails, because agents are like new junior engineers who need guardrails. Enterprise patterns that fell out of favor for being long-winded keep agents in check, and Orosz said, "dead serious," that it may be time to dust off the design patterns books.
Advice: Ship Only What You Can Verify, and Keep Learning
Orosz's first piece of advice is the talk's title. Cap your daily agent usage at what you can review or verify. Verification doesn't have to mean reading every line. Peter Steinberger, creator of OpenClaw, told Orosz he ships code he doesn't read, but builds his own verification systems, thinks at the architecture and module level, and has the AI draw diagrams for him. Either way, don't ship more than you can verify.
Second, tech debt is now cheap to remove: have agents kick off agents to clean it up, and become your team's "chief tech debt remover." Orosz said if you aren't doing this, you aren't using AI efficiently. Third, experiment, since there's no one-size-fits-all. Share approaches with friends. Mitchell Hashimoto, creator of Ghostty and founder of HashiCorp, told Orosz on the podcast in March that he keeps exactly one agent in the background: when he's coding, the agent plans, and when the agent codes, he reviews. He doesn't use multiple agents. Orosz contrasted this with people running five agents, which they personally don't understand how anyone manages, and stressed that Hashimoto cares about high-quality software, not AI for its own sake.
Fourth, spend more time thinking and understanding. Orosz relayed Dax's quip that he used to spend 95% of his time thinking and 5% coding, and with AI now spends 96% thinking and 4% coding. Orosz said most people won't have that luxury but it's a useful frame. Finally, citing Addy Osmani: don't outsource learning. It's too easy to let AI fix a bug while your mental model stays unchanged, trading long-term capacity for present-day speed, and the tools won't stop you. Whenever you use an agent, learn something.
The Job Market and Future-Proofing Your Career
On the job market, Orosz said the global picture looks okay, based on a Pragmatic Engineer deep dive using Indeed data, which they consider reliable. Top tech companies are hiring more than before, though conditions aren't as good as they once were. Orosz cited a 20% change in software engineering postings in the US and UK (introduced as bad news but described as an increase, so the framing in the talk is unclear), decreases of 13% in Germany and 10% in France compared with two years ago, and flat numbers in Canada. They had no Hungarian data but suggested Hungary's close ties to Germany may mean a similar trend. At top US tech companies, AI engineering hiring is booming and now makes up about 10% of all software engineering roles. Orosz defined it as building RAG, evals, and systems that use LLMs.
For future-proofing, Orosz recommended building things on top of AI and LLMs, whether a side project like a podcast recommender or an internal tool at work, and suggested Chip Huyen's AI Engineering as a practical book. They urged engineers not to outsource their thinking, to become product-minded (pointing to the book and their blog post The Product-Minded Engineer), to talk to product managers, and to understand the business. They also recommended becoming a domain expert. At an agriculture company, understand agriculture, because many engineers exist but few have talked to farmers. At an automotive company, talk to mechanical engineers. That expertise makes you valuable during downsizing or a job change.
For engineering leaders, the advice is to be or stay hands-on, "yes, again, otherwise you will be out this time." Orosz said AI makes this easier, since it can explain code and help you contribute. They had heard of companies where five of the top hundred committers are product managers or engineering leaders. Leaders can also help integrate AI at the systems level to remove friction. But they should expect to do less people management, because the business expects it, and those who love people management will either burn out or do less of it. For engineers, that means less career support and probably smaller raises for a while.
Closing: Change This Fast Is Overwhelming, and That's Okay
Orosz closed by noting that Martin Fowler, Grady Booch, and others told them the industry hasn't changed this fast since the 1960s: in about 12 months, AI went mainstream across coding tools. Orosz said feeling overwhelmed is fine and that they sometimes still are, given how fast and unpredictable the change is. Their final suggestion was to pause occasionally, acknowledge that you are keeping up, and ask how to make your work more sustainable, how to produce higher quality, and what you can automate. Then make a change, and repeat.
Good morning, Budapest. It's awesome to be back. Today, I'd like to talk about some people, a few thousand, that are having a terrible week this week. And this is specifically people inside Meta and Instagram. I talk with a lot of people in the industry. I have a lot of friends inside of these companies. I have even more, I guess, contact software engineers who message me to tell me what they're seeing, what's happening. And this week has been the worst in Meta or Facebook history in probably forever.
So what happened is on Monday, we've had the goofiest ever Instagram exploit. It wasn't even an exploit, but it was a security breach. It was an attack. What happened is, I mean, this is from a software engineer writing a book, this is the most goofy thing. So there were two steps to this exploit. Step one is you had to fake your location to a victim. Let's say you wanted to take over Barack Obama's account on Instagram with, I don't know, tens of millions of followers. You fake your location with a VPN to the US and then you went to Meta AI and said, hey, my account has been hacked, could you please send a verification code to this email that I own. And then step two was there was no step two. This was it. Meta AI sent out a code to you and you could take over anyone's account.
This is the first zero-auth password reset, and we're software engineers. You know what this is? This is a bug. But the thing that I couldn't get my head around is I know people at, when it used to be called, Facebook. They have a really strong engineering culture. They have the globe's best automated rollout canary system. They have so many layers of verification. They have really good engineers. They have manual code reviews. And none of it. They have a trust and safety team, for God's sake. Instagram's trust and safety team is closer to 100 engineers whose job is to keep this platform secure.
And, you know, a few things happened. The next day, on Tuesday, Meta's chief information security officer sent an email saying, "I'm out. I'm quitting." This was very interesting because Meta had just kicked off an outage investigation. They call it a SEV inside of Meta. And in the middle of that, and even before it concluded, the chief information security officer is stepping down.
So I asked around. I asked the people I know at Instagram. I happen to know people on Instagram's trust and safety team. Well, turns out they were only on the trust and safety team. So they even shared more details with me. You're hearing this for the first time ever, by the way. It's not in the press. It probably will be. It was AI. Of course it was AI. The thing that caused the issue was AI-written code that was reviewed by AI and not humans at Meta. And I'm thinking to myself, how could this have happened at Meta? I mean, it wasn't just AI. There's more to the story. It was AI maxing, it was layoffs, and it was AI psychosis at Meta.
And what I mean with this is AI maxing. In April I wrote about this new trend called token maxing, which was happening across so many companies, including Meta, including Amazon, including Uber. Engineers were starting to be measured on AI token usage at all these companies and they started to inflate it. They just want to get to the top of the leaderboard. They told the AI, let's do some dumb stuff, and, you know, I get to more tokens, but I don't need the work. And at Meta there was a leaderboard and you could get status like Session Immortal and Token Legend. In April Meta killed this project, but people were burning crazy amounts of things.
Now, AI usage inside of Meta was part of performance evaluation. It wasn't officially made up, but people inside of Meta are smart. If you had a low token count, you know, that's not a great signal. So people just started to inflate their token count. So they started to use AI for anything and everything. Write it by hand? Nah, why do it? Ask the AI. Read the documentation? Nah, let me use the AI to read it for me so it can just burn a bunch of tokens. This is the craziest thing that's happening, but, you know, inside Meta AI is free. And again, these people want to have higher bonuses. And they just use AI for everything. And yeah, the code that caused this SEV was also AI generated. Of course, they used AI to review it as well. They used it to triple review it, etc.
The second part to this thing was layoffs at Meta. Meta told 10% of staff, 8,000 people, that they were laid off on the 20th of May, but Meta told people, or the press told everyone, that the layoffs were coming a month before. So what people were doing is, as they were thinking, oh, am I going to be laid off, all of them started to use more AI because they didn't want their token numbers to be down, because they didn't want to be fired for not using enough tokens. You see where this is going, right? And they were not really busy, you know, doing their work. They were just worried about, all right, let me get this inflated. So inside of the trust and safety team, people were not thinking about trust and safety; they were thinking about token maxing.
And finally, I was wondering if I should call this AI psychosis, because psychosis is a very serious psychological condition, and that's why I put it in brackets. But I'll show you why I chose this name. Instagram had a trust and safety organization that was built up over like seven or eight years. A really good team, mostly based in London. 40% of this team before the 20th of May was reassigned to do manual data labeling.
They were told on Thursday that starting on Monday, you are no longer working on this team. You are moving to this new team in Alexandr Wang's org, ex-Scale AI, and you will be doing AI data labeling, which means you get these tasks, it's a GitHub pull request, you need to review it, you need to add some tests, you need to add some feedback, and then you do the next, and then you add some tests and you add some feedback and you do the next. These highly skilled people were not given a choice. Now, inside Meta, until now every engineer was treated like royalty; they were given a choice. They were not given a choice. So 40% of the organization just, boom, gone. There's closer to 5,000 developers inside of Meta doing manual AI labeling. And there's a running joke inside of Meta that this data labeling org is bigger than OpenAI. Meta clearly wants to build this amazing AI model.
Oh, and after the layoffs and after the reassignments, most teams are less than half the size. Some don't have on-call coverage anymore, which means that in some services, there's just no one picking up on call. This again has never happened inside of Meta. And this is what I mean by AI psychosis. This is fully self-inflicted. This is coming from Mark Zuckerberg. This is coming from Alexandr Wang. This is coming from the top. They're saying we don't care. It's so important for us to build this model that we will risk our business, and we don't care if, you know, we get hacked or something like that.
Morale is as low as it has been at Meta. I've seen low morale. This is way worse than the 2022, 2023 layoffs. Oh, and yeah, if this wasn't enough, in the US, they're recording your screen. They're recording all your keystrokes to train an AI. So it's the easiest time to recruit from Meta right now. And the engineers that I talked to, they just feel super let down. Meta used to treat engineers like royalty: salary, compensation. You could choose your team. It was a good world. And the CEO, Mark Zuckerberg, is a software engineer. He wrote a lot of Facebook's code. And they feel we don't matter anymore. We're tools. We've been thrown away.
A lot of people who have not been fired have been given large retainer bonuses: you stay, we're giving you money. They are still interviewing, and they told me they're interviewing not because of the layoffs, because they know they can find a job at Meta. You know, these are super highly paid people. They can get not-as-highly-paid jobs. They're in demand. But they said, "I don't know if I'm going to be assigned to data labeling. And as a professional, I did not sign up to become a manual data labeler. And a bunch of my colleagues are doing that and interviewing." So Meta is destroying the engineering organization that they've built up over 22 years. And I think this might be the end of the incredibly strong Facebook engineering culture that I know and I've learned to actually love, even though I've never worked inside of Facebook, and so many of my friends have. So this is because of AI.
Now, not all companies are like Meta. This is pretty extreme, but it's happening. It's happening right now. These are facts. But the industry is having a pretty interesting time. And today, I want to talk about this. I want to talk about how everything has changed in the past six months when it comes to software engineering, or, well, at least coding. I'm going to give you a tour of the tech industry, of what other tech companies are doing. I'll give you a brief tour, because this is what my head is in day in, day out. And again, I talk with a lot of these people. I visit these companies. I'm friends with a bunch of them. And then I'll share a few trends that are happening across the tech industry. And then I'll close with advice to software engineers and engineering leaders on how we can navigate, prepare, do the best that we can, and also just, you know, come out of this whole thing stronger.
So everything has changed in these past six months. This is a pretty dramatic thing to say, but it has changed. So DHH, David Heinemeier Hansson, creator of Ruby on Rails, was on my podcast in February, and he told me, and he actually wrote this on Twitter, that just in summer 2025, he spoke with Lex Fridman in October and he said that AI was not writing any of his code directly. But part of his resistance was that the models were not good enough, and it has now flipped by February. Most of his code is being written by AI. Now, this is a person who, you know, is big into software craftsmanship. He is not paid by any lab, but he decided the models are now good enough. They write better code than I can. He actually told me this on the podcast. So I listen to people like DHH in this sense.
Simon Willison is one of the most quoted people on Hacker News. He is an independent software engineer. I love Simon. He's also a friend. He created Django, and he writes this really good daily newsletter, pretty much, where he just experiments, he builds open source and tries out all the models. And he said that the models released in November 2025, specifically Opus 4.6 and GPT 5.4, have elevated agents to being genuinely useful. We've had the six months to get used to this idea now. So no wonder that companies are now starting to spend big money on this. So he's also saying that it has changed.
I got some data from some of our partners. My friends at Linear shared this data, never shared before, on how teams are now using agents to ship more code. They are comparing teams that are on Linear and not using AI agents with teams that are, and by now the teams that are using agents are shipping five times as much code. We'll talk about quality later, but this is a massive 5x increase. I mean, we've probably had a five times increase in like 20 years before, and this is in less than two years.
Friends at Cursor have shared details on how devs using Cursor are changing the lines of code they produce in a year, and in a year it has gone up by almost two and a half times, from 3,400 lines of code to more than 8,000 lines of code, just on Cursor. So, you know, we're seeing this acceleration. The size of PRs, also from Cursor, is up by 3x. So if you combine those two, that's six times as much code. And, you know, a lot of you, I see, are smiling. We know that there's six times as many bugs. Yeah, again, we're going to get there.
And also data from Cursor: the percentage of devs using Cursor who are accepting changes from the AI without any manual review is massively up in January. This is when Opus 4.7 came out, GPT 5.5 came out, and what a lot of companies realized is that Claude Code, Cursor, Codex are actually really, really useful, and they're starting to trust them. And again, remember when I told you about Meta merging to production without human review? Yeah, that was somewhere there.
So, what are tech companies doing right now? Let me give you a bit of a tour of the industry. So at Anthropic, I visited their offices last fall and I talked with Boris Cherny, the creator of Claude Code, just in February. And here's what they're doing. Boris specifically runs five parallel agents on his laptop all the time. He ships 20 to 30 pull requests a day, and this is on top of leading all of Claude Code. He's very much a hands-on leader. He told me that PRDs, writing documents to plan, are dead. They're using prototypes across Anthropic to replace them. Today 100% of Claude Code is generated by Claude Code. Inside Anthropic it's not 100%, but it's closer to 70 to 90%, and there's no target. This is just people using it. Again, this is Anthropic; we shouldn't be too surprised.
And then the company built Claude Cowork in only 10 days. And it's become a massive commercial success for them. It generates so much revenue. In fact, I have some sources inside of Microsoft. Microsoft tried to build a Claude Cowork, because Claude Cowork is really good for Excel and Windows, to use it on the machine. Microsoft still doesn't have an answer two and a half months later. I heard that Satya Nadella gave this team a deadline of a month to build it and they couldn't build it. So don't forget that there are differences between companies. Anthropic is accelerated by AI. Some companies are kind of held back despite AI, like Microsoft, and again, we'll talk a little bit about that as well.
OpenAI. OpenAI also, a friend from the Pragmatic Summit in San Francisco in February. That's us on stage, with him and the Codex team. We talked about a bunch of stuff, and he told me some interesting stuff. Inside of OpenAI they have an internal version of the ChatGPT app and they have a fix-it button. You can literally just take a screenshot and say fix this bug, and it goes to Codex. It generates a pull request and an engineer can merge it. In fact, even a non-engineer can merge it, and there are safety nets there. AI code review obviously is everywhere. They have multiple layers of it. They have tiered versions. There's some code that can go into production with just AI code review, and there's some code, the critical path, that humans need to review, engineers need to review.
Most devs obviously run several agents. There's this joke that when you're walking around, engineers are bringing their laptop and it's slightly open. It's slightly open so the local agent can still keep running. And I did a video interview with one of the OpenAI folks, and I was just jokingly asking, "Oh, so throughout this interview, did you have agents running? Did you have an agent running?" He's like, "I didn't have an agent running. I had five." And I was like, "Oh, okay." It's common for people to go into meetings and their agents are running. They're thinking about agents. They keep them on track. But again, these are the most AI-pilled people in the industry. And of course, you know, they greatly believe in all this. They all talk about AGI and when it's coming, not if it's coming.
But inside of OpenAI, most people don't really write code. This changed in October. There were devs who wrote 30% of their code by hand and 70% with AI, but I think that 30% by hand is just slowly going away. The Codex team obviously writes it all with Codex.
they're telling me that taste, knowing what to build is becoming pretty important inside the company. Codex also improves itself, as a fun fact. It tests itself all the time. It runs all the tests overnight. They kick it off because most of the team is in San Francisco, so it's one time zone. They have Codex run itself and look for ways to improve itself. And by the morning it comes up with improvement suggestions which they either accept or reject. And when they have meetings and debugging sessions, when they start the meeting they have voice notes that they send to Codex as it goes, and it comes back by the middle or end of the meeting with results. It sounds like science fiction, but again, that's how they're working inside.
Cursor: I visited their office in October in San Francisco. They have you take off your shoes, and sometimes it's a mess, sometimes it's super organized. There's like a sorting algorithm invisibly happening inside of their office. It's really interesting slash cool. But they're a very nice group of folks. They have gone all in on agents as of January as well. It used to be all tabs and the editor. They're kind of moving on to agents. They still have the old experience, but it's increasing the old one. They built their own coding model. They're one of the only companies outside of OpenAI and Anthropic who have a really good coding model. I have no affiliation with Cursor, but their Composer model is cheap, which is going to be important, as I'll talk about.
And they operate tens of thousands of NVIDIA GPUs in massive data centers. They're leasing it from Azure, AWS and so on. And most of their inference used to be — inference is generating the response — it used to be their biggest cost, but now they're also training their models. So they're kind of turning into this mini AI lab. And of course now SpaceX is about to purchase them, or not, or who knows, but it seems it's going to happen.
Also, everyone at Cursor is technical. This is Lee Robinson, developer relations at Cursor. He wrote, with Cursor, he migrated all of Cursor's sites to a different CMS, and of course he's sharing how much it cost to show that it's very economical, but this is not a software engineer by job, and everyone at Cursor is like that. So these labs are — everyone goes there.
Google, briefly. Everything is custom at Google, everything including their IDE. Google's internal IDE is called Cider. They have strange names for everything. It used to be a web-based tool; now it's a Visual Studio fork. They have a thing called Jetski, which is Antigravity but the internal version, which is integrated with their monorepo Piper and all of their other internal systems. They have Critique, a code review tool — again, they don't use GitHub, everything is custom inside of Google. AI is of course integrated; Gemini is integrated nicely in there. They have Code Search, which is the Sourcegraph for the rest of the world. In fact, Sourcegraph got inspired by it. Google has some of the best code search inside of Google; they don't make it available outside. And Google has so many internal systems: Borg, which is their version of Kubernetes, Monarch, which is their version of Datadog, many more, Piper, their version of a monorepo. AI is integrated into all of these things and it's all integrated together really nicely, so inside it's a really good experience. The only problem inside of Google is Gemini is just not as good as Opus or GPT 5.5, and inside of Google, whenever engineers can use Claude Code they do, but they can only use it inside the Gemini org, which means that Google doesn't have as good adoption of AI as some of the other companies. Kind of weird, but they're working on it. The CEO knows, he admitted it. They want to get a better model.
And finally, Meta. They want to build their state-of-the-art AI model. Everything is about this. They do have an internal tool, Metamate. That's their AI tool for coding. They have this thing called trajectories. When you see GitHub commits inside of Meta, you see the exact prompt that people used. They rolled it out in December and people in Meta got upset because no one told them this would be public, and you could see staff engineers saying, "Can you write me a for loop?" and it was all public and everyone could see it. So some people inside of Meta — I talked with this dev and he said, "I started to write my chats with the Meta AI, the code generation, in Polish because fewer people can read it now."
Okay, but right now at Meta they have bigger things to worry about: this forced reassignment to build an AI model, forced tracking of everything. It's clear Mark Zuckerberg wants Meta to have a model that's better than Opus 4.8. I think either he's going to get it in the next few months, or a lot of Meta is going to be, not disbanded, but very, very demotivated.
Uber, my old company. I talked with them in detail. I have a deep dive on The Pragmatic Engineer if you're interested in learning more of these details. They built so much in-house tooling, and a lot of companies do this, but I'm just going to quickly show you how much in-house AI tooling a company like Uber built. Uber has about 3,000 engineers, so just keep that in mind. They have a developer experience team, who is now pretty much an AI experience team, of about 20 people or so.
So they built an internal MCP gateway. Pretty clear: you can discover, register, do all sorts of jazz. They built an Uber Agent Builder, which is a no-code way to build agents for the rest of the business. They have an Uber Agent Studio where you can drag and put together your agents. Again, OpenAI has something like this publicly, that's OpenAI Agent Builder, but it's for the nontechnical folks. They have Uber Agent Builder Registry — so Uber is 3,000 engineers, but 20,000 other people. Those 20,000 people use this thing. They create this stuff, they plug it up, and engineers built this for them.
Uber has an AIFX CLI. I'm just going to call this the Claude Code for Uber, pretty much. They built it themselves, integrated with all their systems, using all the different models, etc. They have Uber Minion, which is running background agents at scale. So again, this is a similar thing as Cursor background agents, except it's integrated into Uber's monorepo and experimentation system called Morpheus and all of the other jazz really nicely, and it works a lot better. So even though devs can use Claude Code, they will use Minions because it just works better and faster. For example, Uber Minions, when you give it a prompt, will analyze it and give you a suggestion that these prompts could work better, get results faster, cheaper, etc. Claude Code doesn't have this yet.
They have Uber Code Inbox. People are getting so many pull requests and code reviews, which are now mostly AI reviewing code, that they're creating a system to show: this one needs your attention, this is important, focus on these things. So when people get into work, they start going through these things. They're trying to make code review a bit more fun. They have something called smart assignments, where there are SLAs: if this person doesn't respond in like a day, it goes to the next one. It's a bit like on-call tooling, again all custom. They have risk profiles. They will try to identify that this code change looks freaking risky, you need to look at this closer. And they have uReview, which is the CodeRabbit or the Sonar for Uber internally, again all custom work. So they build all this: MCP, agent builder, CLI, Minions, etc.
And the other large tech companies are doing the same. I'm not going to run you through all this. Stripe has Minions, Toolshed, Blueprints, Z boxes. Ramp has Inspect, Glass, Dojo, Sensei. Sensei is a funny one. Shopify: Sidekick, LLM proxy, Dev MCP server. Airbnb: One everything, Catalyst and so on. They all build their own stuff. They have a dedicated infra org building all of these at all these companies. So if you thought you're pretty cool for integrating an AI agent into Slack, you are pretty cool, but this is next level.
I talked with a bunch of startups, and I'm not going to go through all of them, but the general trends I see there are kind of the usual: agents are doing coding, doing code review. There's a bunch of creativity, mostly about Slack. People tag Slack. I saw a startup recently that raised $70 million in Series B. They just told the agent, "Fix all bugs in the codebase," haha, and everyone's laughing in Slack. And then the agent came back like, "Oh, I actually found four critical authentication issues where your back door was wide open." And people were like, okay. I mean, that's what startups are. They didn't know how exposed their house was. They're usually plugging in the AI agents, integrating them, and some of them are having fun vibe coding SaaS. I think it's just engineers having fun. I don't think it's really a business thing, but I only see this inside of startups, not really inside of big companies.
And inside traditional companies — this is the most interesting thing — it's all the same. I mean, not at the level of Uber. They don't have dev platform teams, but they are not really lagging behind. For example, Cisco rolled out Codex to 18,000 engineers back in January, when Codex was pretty small, and they're doing a bunch of complex migrations. JPMorgan Chase built a multi-agent framework, which is a fancy way of saying that it just uses multiple specialized agents to label customer interaction data. They use evals, judge-based aggregates. It's kind of cool stuff, even inside of these companies.
So this is what's going on inside. Now I want to give you some of the trends that I see crosscutting everywhere, or mostly everywhere.
One of the big things comes from Laura Tacho. This is me at the Pragmatic Summit with Laura and with Martin Fowler in San Francisco. I messaged her actually last night and she replied this morning. I was asking, "What do you see, Laura?" because she was CTO at DX; she's now heading up pretty much developer experience at AWS. And she said that many organizations get stuck: they see individuals doing great, but the team output is not there. And she said it's because they are thinking about AI as a productivity tool for individuals, and she calls it the individual speed-up juice: things like email summaries, Slack automations, even code generation.
However, the companies that are moving faster and are seeing the results for teams are doing something different. They begin with a business outcome. For example: I want to deploy to production faster, or I want to push more features out with the same quality, or I want to improve quality.
Spotify is a very good example. We don't hear too much about Spotify, but I talked with their CTO about a month and a half ago. We had lunch, and he told me that their bar for using AI is that the quality needs to stay the same. So they're not seeing a huge increase in output, but they have built a lot of internal tools to check for quality, and they're slowing down the rollouts of AI versus what Meta is doing or whatever. And again, that was their goal at Spotify.
And Laura was saying that you want to build an agentic system that reduces handoffs, that makes it easier to find information, and removes friction while maintaining quality. That last part is very important. Few companies do that. And maturity comes from applying AI to the system and not the individual. A lot of people are focusing on the individual, and that's why we're not seeing it. And she created this mental model. She was saying that when you have AI usage that is either individual or team level, and decision-making that is either simple automation or agentic systems, most companies are in this bottom left corner, where you have individuals doing simple automation. Where most companies want to be is where they have team-level agentic systems. But to get there, you need to do what I've shown you Uber doing. You need to build a lot of systems that integrate. You need to iterate on this. It takes time. It takes a massive investment. You're not going to be able to buy Claude Code or Cursor or whatever vendor tells you that it does it, because you need to build it into your system with your engineers. That's what Uber is doing, for sure.
Now, another trend: token maxing and tooling addiction. Some of you might be doing it, some of you might not. It's going out of style, by the way, just as I'm talking. There is just a big pressure to look productive and to not have a low token count, especially inside of US tech companies that don't really care about budget until they do. But right now some of them still don't. It's ending. Token maxing is when you're just burning all these tokens without value shipped, on purpose. And again, I've talked to people, this happens at Meta, Amazon, even Microsoft, everywhere where they have internal leaderboards. Microsoft still has it. I don't know why they're not shutting it down. They should listen to me.
Also, the pricing of these tools feels a bit addictive. You buy the $10 plan or the $20 individual plan, and then you run into a limit, a generous limit, but then you run into a limit and you're like, "Let me buy the $100 plan or the $200 plan." And once you buy it, if you're buying it for yourself, you now feel pressure that you're not using your allowance. So you start to use it more. And next thing you know, you went over and you're now on API pricing. And also, with every prompt, once you start using the AI agent, the first few months it's a bit like gambling. Some people get sucked into it, really like gambling. It's just one more prompt, one more prompt. People are not sleeping that well. You're waking up and thinking about your agents. If you're paying out of pocket, you feel AI being wasted. It's weird. It's addictive.
Another trend is middle management. Managers are just being cut: either laid off, or, inside of Meta, reassigned to individual contributor roles, or being told you need to be hands-on, meaning you need to manage less and do more work. There's just a flattening happening. And whenever a manager is fired or laid off, it's said, "Oh, it's because of AI," whatever. It doesn't help. But the interesting thing about this is: what happens if we have less middle management? I mean, it's popular to hate on middle management, on managers, senior managers and directors. Top-level management is C-level, the CTO, and middle management is everything in between, maybe until frontline management, engineering management. And usually we don't know what directors do, or if they're necessary.
However, in my experience, good middle management, good directors, good senior engineering managers are very technical. They could be hands-on, but they choose not to. But they listen, they see what's happening, they pay attention, and they make small changes. "There are a lot of outages we're having right now." Software engineers will just pile on and do nothing. They will stop and be like, "Okay, let's create a task team. Let's build this system. I will pull you off these teams and we'll make our engineering culture better." Good engineering management improves engineering culture, and a lot of companies are getting rid of engine
management, or bad management, and engineering culture will go down. This is a fact as far as I'm concerned.
Another interesting trend at the same time: CEOs and CTOs are back to coding. Guillermo Rauch, founder and CEO of Vercel. I had a lunch with him at one of the investor events in February as well. He was writing recently that he is seeing so many CEOs and CTOs are back to coding with a fury, with all this enthusiasm, and he has public company CEOs DMing him saying, hey, we're using Vercel or Claude Code and, you know, I'm doing it, I'm so excited again. And this is all the time while we're having less middle management. Now imagine having less middle management to protect engineers, and the CEO and CTO are coding, vibe coding, and they think it's complete, but you know, it's not really complete.
A mega trend that is happening, and I noticed this a week, two weeks ago. So I wrote about it a week ago, and then today, hold on, and then today I see Sam Altman, this is just from this morning, saying that he is noticing that AI budgets are seemingly becoming a huge issue for some companies, and something that has come up and something that has never happened before. And I was pinging people at OpenAI like, does he read my newsletter? Because I wrote about this last week for subscribers, and someone at OpenAI said someone posted it into Slack and Sam read it. But it's happening, and this is coming out of the blue. There's this joke going around as of yesterday on Reddit saying, hey, oh baby, I see $15,000 are gone from your shared account. Is this what I think it is? Engagement ring?
Yeah, I feel for that guy. He's soon going to be single, assuming it's not a joke. But it's happening and it's getting worse. Anthropic has turned on API pricing for enterprise customers, meaning anyone who's not a startup or an individual is not getting discounts. GitHub Copilot turned it on just two days ago, on the 1st of June, and people are pissed because they have burned through their usual budget of, let's say, $200 or however much it was in three days, that used to take a month. And this is hitting everyone right now. Everyone will be paying a lot more.
Now, Uber is an interesting case again, because in March their CTO said that they have burned through the whole budget for the year with AI costs, and we were wondering what they're going to do, but we now know they are now setting a cap of $1,500, $1,500 per month per engineer on AI, and if you hit that you're going to use the free models. And I've been doing research; this is what a lot of companies are doing. A lot of companies are not doing this much, some are doing $200 and then you're going to use 0x models on GitHub Copilot, and now engineers want to do it. But this is a very, very fresh trend. Costs are... it's ridiculous when it's as much as an engineer, and no one wants to pay that, no matter what the AI labs say.
And finally, some trends across the software craft. We've talked about business trends and AI trends, but what is happening to software engineering and the craft, the conference that we're here at? One is a huge drop in quality everywhere. This one comes from yours truly, that's my account. I was so pissed off at claude.ai, their flagship website, for about a month. For a month, every time I went to the website, this is the website itself. I did a screen recording after I got pissed off enough, because it kept happening and no one was fixing it.
You went to the main website, claude.ai, and I immediately start typing my query. And here I'm starting to type, "How can I do this?" And as soon as I type "How can", there's a refresh. Now, there's a React lifecycle component happening here where the page finally refreshes and it loses all that I've typed before. And maybe I'm old school, I use the website so much, it just kept happening and happening, and finally I tweeted about it saying, how on earth does Anthropic... oh, and I'm a paid user, I'm not a free user. There's millions of people hitting this every single day, and Anthropic doesn't care, and they're building AGI. So I tweeted about it, and the product manager on the team said, "Oh, great feedback. I dug into this, it will be fixed." This is the short way of saying, "Oh, thanks. We had no clue that millions of people every day are doing this. Oh, and we are not even dogfooding our own stuff." And this was there for a month. Oh, and we're the fastest moving and biggest and most profitable company, but we don't... A bank does so much better in this sense. There's not these... I mean, we can argue if they fix it that quickly, and they did fix it eventually, but this is Anthropic. And it's not just Anthropic.
OpenAI bragged about how they built this amazing Agent Builder that is similar to Uber's internal agent builder in only six weeks with one engineer, with, you know, Codex. Amazing. Great. Quality is terrible. People on launch tried to use it and they kept running into so many issues. Their forum is full of comments which are unresolved. OpenAI did not come back and fix it. This is from three months after launch, someone saying, "I was bullish on Agent Builder when it came out, but, for example, P0-type bugs are not getting fixed or take ages. It just seems like abandonware." So I mean, was it worth it for them, building this thing and then just forgetting about it? And AI clearly didn't help build higher quality software. It's faster, but it's just...
Amazon, AWS: an engineer allowed the internal Kiro AI coding tool to make certain changes, and the agent opted to delete and recreate an environment inside of Amazon, causing a massive outage. Amazon had AI bugs that happened because of AI-generated code, where Amazon's store, their .com flagship website, part of it went down. This never happened with Amazon. Same thing as it never happened with Meta, and this is over-reliance on AI, or not caring about quality. Amazon has made this change that it now requires a senior engineer to review any AI-generated change, because they realized the junior engineers will just say "looks good to me" and it causes an issue.
OpenCode is the leading AI harness. They're like the Claude Code for open source, and they use all different models. Dax Raad is the founder, and I love Dax because he's super honest. They are building a super popular AI tool. They have almost a million daily active users. They're growing; they've grown 10x in the last four or five months. And he's kind of skeptical of AI hype, but this is a guy who's built developer tools. I love Dax, he's really authentic. On my podcast just last week he told me, "We're shipping way more hacks where we should have first rethought the whole system from the ground up, redesigned it to make it more flexible. So I think our judgment," meaning the OpenCode team's judgment, "is off."
And he was also saying how, you know, we're in the AI coding tool space, but you know what's not happening? No competitor is beating us because they're using AI better than we do. And he said that, frankly, I don't think we're using AI that well. We're actually telling ourselves to use less AI, and there's no competitor that is beating us because they're doing it faster. In fact, they're kind of winning, because they're still one of the most quality harnesses, because they're slowing down. Do you know what is a contradiction? A CEO and founder of an AI company saying we need to use a bit less AI. And he actually told me, we need to do more thinking, we should build fewer things and build the things that matter.
I'm paying attention to him. I'll send that to Dax.
And another trend related to this is just everything is broken. GitHub is such a prime example. This was two weeks ago: all your pull requests were gone on GitHub for about 8 to 12 hours. There's an alternative GitHub uptime tracker, I think you just have to search for "the missing GitHub status", something like that, which tracks all outages that they report, and based on this estimate they don't even have one nine, which means some part of GitHub is down 10% of the time, which is absolutely unserious. But this is a serious company. I talked with the GitHub team, I talked with their COO, and they gave me data that they didn't give anyone else, because they published graphs without the numbers, but they gave me the numbers, and they told me it's because of the load. Now, the load is this: it is a 3x load increase over two years' time. And they were like, oh, you know, this is a huge load increase, we could have never prepared for that. And I'm saying: really? That's it? This is bringing GitHub down to nines? I do not buy this. Maybe there's other things, but something was really broken inside of GitHub. I'm not going to say this is AI-generated code, but if GitHub cannot deal with a 3x increase over two years, and sure, this will be a 5x increase later, you're doing something wrong, guys. Other startups pick up this load laughing. And there's details, GitHub has a Ruby on Rails monolith and so on and so forth, but yeah, it's just breaking.
Mario Zechner, the creator of Pi, which is what powers OpenCode, this is the Austrian... it's with Armin Ronacher, the Austrian AI mafia, who were on my podcast. He told me, "It just feels software has become a brittle mess everywhere. 98% uptime feels like the norm on most services. User interfaces have the weirdest bugs." I showed you one on Claude, but it's everywhere. And he says, "I give you that it's been the case for longer than agents exist." We've always had it, "but it feels to be accelerating everywhere." You see this. I even saw it with Magyar Telekom the other day. I don't think it was AI-generated, because I don't think they use AI, but yeah, I had a big software issue with them and I needed to call customer support.
One more trend is slop buries the software engineers who still care. Here's what's happening. There's a lot more pull requests. There's just a lot more code, and a lot more of it is AI-generated. Most developers inside a company have review fatigue, and they see it's AI-generated, their AI review went through, and they said it looks good to me, LGTM, or, you know, I'm not sure how, but they just do a thumbs up and they never reviewed it. There are a few developers who do review it. Hands up if you actually still review code properly. Hands up if you give it an honest shot.
Yeah. But there are many of you who still try, and you still catch the bugs, and you still push back, and you still see that the agent has duplicated code, or, well, the developer is the agent. You push back, and they are being overwhelmed. They are being burnt out. They are being fed up. They are feeling that they're not rewarded. Oh, and when it comes to performance review time, they're not going to be rewarded. They're not seen as the ones pushing out all the features. So some of them are burnt out and some of them just quit. Dax told me that at OpenCode they are hiring a bunch of these people who are leaving their companies because they're just burnt out being the sole person still keeping things alive, and no one cares. Engineering management is gone. They've either let them go or they're now less hands-on. So there's no one left to care.
Finally, I talked with Kent Beck. He'll be the keynote speaker tomorrow, and he summarized this really well. Kent is amazing at summarizing findings. He said, "We're accumulating code faster than we accumulate trust." He said that with code, you need to trust it, you need to understand it. We don't have time to do that right now.
AI also amplifies software engineering experience. So seniors gain the most; judgment is rewarded. And we see this everywhere. Hillel Wayne, he'll be a speaker tomorrow, but he was telling me how some people are saying, oh, AI will help with formal verification, with TLA+, a very complicated language. He'll show you a demo tomorrow. And he said the only people who have been successful with AI generating TLA+ specifications that work are TLA+ specification experts who, in the prompt, gave the exact specification of what kind of prompt to generate. Everyone else, good luck with that. And this is true for software. If you're a junior engineer, if you've never built a mobile application, you can prompt a native iOS app, you can prompt the agent, it'll build something, but, you know, it's not going to be maintainable.
Old patterns seem to be coming back. Dax told me how domain-driven design and verbose guardrails, they're using this at OpenCode all the time, because agents are the new junior engineers. You can start off a lot of them, but these junior engineers, I mean, if you think of it like that, they need a lot of guardrails. And these boring enterprise patterns used to become unpopular because they're long-winded, you have to explain, you have to type out, but they keep agents in check. So it might be time to dust off some of these books and start to use design patterns again. I'm actually dead serious about this.
So this is where we are. It's just a lot of change, all sorts of things. It's confusing. I'll leave you with a little advice. One is the title of the talk: slow down to speed up.
My suggestion is to cap your daily agent usage to what you can either review or verify. You might not need to read the code. Peter Steinberger, creator of OpenClaw, I did a podcast with him. He said that he ships code that he does not read, but he builds his own verification systems. He thinks in architecture. He always looks at the module. He has the AI draw diagrams for him. So verify: do not ship more than you can verify, or this might mean reviewing. By the way, as well, tech debt is now very cheap to remove. We're talking about how it's built up. Have the AI kick off the AI to remove the tech debt. Be the chief tech debt remover on your team. You will feel better for it, and it's much easier. Forget that it's hard to remove it. It's not. And if you're not removing it, you're not using AI efficiently for yourself.
Experiment with different usage of AI agents, because there's no one-size-fits-all, and, you know, talk with friends about how they're using it. Get ideas, tell them what you're doing. And, you know, just spend more time thinking and understanding. This is what Dax, the creator of OpenCode, says: that he used to spend 95% of time thinking and 5% of time coding. And he said, "Yeah, AI is cool, because now I can spend 25% less time coding. So I spend 96% of time thinking and 4% of time coding." You're not going to have the luxury, but, you know, it's a good way to think about it, and you just spend time thinking and understanding.
One thing about working in a different way: Mitchell Hashimoto is the creator of Ghostty, founder of HashiCorp, a really nice guy. He was also on The Pragmatic Engineer podcast in March, and he came up with this rule for building software. He is not an AI maxi, but he likes to be productive. He has one agent in the background always doing something. He said, "If I'm coding, I want an agent planning. If they're coding, I want to be reviewing." And he always has just one extra agent. He does not use multi-agents. But this is what I mean by experiment. Some people use five agents and they can manage; I don't know how, personally. Mitchell found that he can use only one agent. He has this buddy that he has all the time, and he said it works for him, for now at least. So just experiment, try different ways of doing things. You don't need to go overboard. He's a very productive engineer. He does not care about AI; he cares about writing high-quality software.
From Addy Osmani, I met him two years ago back at Google, and also a friend. He said one of the best things: don't outsource learning. It's just too easy to let the AI write code while you skip all of the learning. The bug gets fixed, but your mental model does not. We're trading off your capacity for present-
day speed, and the tools don't force us. So, whenever you use these AI agents, do not skip the learning. Understand, learn something when you use an AI agent. And this is what I mean: you don't need to review all the code, but you need to learn something from it, or build a system, have it build something for you.
When I look around on the job market, I have some good news. The job market seems okay globally. We did a deep dive in The Pragmatic Engineer, and you'll have access to that deep dive, the full one, very soon. The top tech companies are hiring more than they have before. So it's going up. It's not as good as before. Now the bad news is that in the US and in the UK, we're seeing 20% increase in software engineering. This is not AI engineering. This is software engineering. In Germany and France, we're seeing 13% and 10% decrease, which is not terrible, but it's not great. This is from two years ago, and in Canada, it's flat. I don't have data deals from Hungary, but as we know, Hungary is very much tied, as we know on the press, to Germany a lot. So, a similar trend might be happening. This is data from Indeed. It's a pretty reliable data source, and I trust them.
Now on the job market in the US and US tech companies, the top tech companies, AI engineering is an absolute blast in hiring. So AI engineering is part of software engineering; it's now taking up about 10% of all software engineering and is going up. So AI engineering means you're building RAG, you're building evals, you're building systems that are doing something with LLMs.
Future-proofing your career, my personal advice: build things that build on top of AI and LLMs, because you will get hands-on with RAG. AI Engineering, the book, by Chip Huyen. It's a very practical book. And try to build something, either on the side, build your own podcast recommendation system or whatever, or at work. Show off to your colleagues and even your managers, and your colleagues will be happy to see, oh, really cool, like you just built an internal tool to do XYZ.
Don't outsource your thinking to AI. Try to think more on product, understand the business. There's a book called The Product-Minded Engineer. I have a blog post called The Product-Minded Engineer. Talk to product managers, and just understand how the business works. It is more important.
And try to become a domain expert. Try to become this industry insider. If you are working in an agriculture company, understand agriculture, because there's a lot of software engineers but there's very few who have talked with farmers. If you're working at an automotive company, talk with the mechanical engineers as well. Again, if you build that domain expertise outside of software engineering, you will be in demand the next time your company either does downsizing or you want to move elsewhere.
If you are engineering leaders, my advice to you: you need to be hands-on or you need to stay hands-on. Some of you will think, not again. Yes, again, otherwise you will be out this time. And it is easier to do this with AI. You can turn to AI to explain stuff for you. You can start to contribute stuff. I'm hearing of companies where, of the top hundred committers, five are product managers and so on, or engineering leaders. And also you can just help integrate AI at the systems level. That removes friction.
Except you will be doing less people management. The business expects you to do less of it. If you love doing people management, either you burn yourself out or you will do less of it. And if you're an engineer, if you're a developer, you will get less career support. Also, we're probably going to see less pay rises and some of those things for a while. But again, we will just have less management, with all the good and all the bad parts of it.
Finally, I'll close with this. I talked with Martin Fowler, I talked with Grady Booch, I talked with all these people. They all said change has never been this fast in the software industry since the 60s, easily. In 12 months, we've had AI go mainstream across coding tools. If you are overwhelmed, absolutely okay. A lot of us are overwhelmed. I was overwhelmed. Sometimes I still am overwhelmed with how fast this changes and how unpredictable it is. But from time to time, pat yourself on the back. You are keeping up. It's hard.
Look around. Just stop for a little bit and make a change. How can I make this more sustainable? How can I produce more quality? How can I automate some of these things? And then rinse and repeat.
Gergely Orosz, everybody.
Article published
