Slow Down to Speed Up: Gergely Orosz on What AI Is Actually Doing to Software Engineering

Open on YouTube ↗
Overview

In this keynote at Craft Conference 2026 in Budapest, Gergely Orosz, author of The Pragmatic Engineer, asked whether AI coding agents are really making teams and companies more productive. Orosz relies on sources inside many large tech companies, and argued that coding has changed dramatically in the past six months, but that the results so far include collapsing quality, runaway costs, burnout, and in one case a self-inflicted crisis at Meta. Their advice, captured in the title, is to slow down: ship only what you can verify, keep learning, and put AI to work at the level of the system, not just the individual.

29 min read

The Instagram "Exploit" That Wasn't Really an Exploit

Orosz opened with events from the week of the talk, which they called the worst in Meta's history. On Monday, Instagram suffered what they called "the goofiest ever" security breach. It had two steps. First, an attacker used a VPN to fake a location matching the victim, for example the US if the target was Barack Obama's account. Second, the attacker told Meta AI that their account had been hacked and asked it to send a verification code to an email address they controlled. Meta AI sent the code, and the attacker could take over the account. Orosz described it as effectively the first "zero-auth password reset," and said that to engineers it is simply a bug.

What puzzled Orosz was how this could happen at a company with such strong engineering culture. They listed what Meta has: what they consider the world's best automated rollout and canary system, many layers of verification, manual code review, and an Instagram trust and safety organization of close to 100 engineers. None of it stopped the bug. On Tuesday, Meta's chief information security officer emailed staff to say they were quitting, in the middle of an incident investigation (a "SEV" in Meta terminology) that had not yet concluded.

Orosz said they had talked to people on Instagram's trust and safety team, and that the details had not yet been reported publicly. According to those sources, the faulty code was written by AI and reviewed by AI, not by humans. Orosz argued that AI alone did not explain it. They named three contributing factors: token maxing, layoffs, and what they called, with the caveat that psychosis is a serious condition, "AI psychosis" coming from Meta's leadership.

Token Maxing, Layoffs, and Forced Reassignment

Orosz had written in April about "token maxing," a trend at companies including Meta, Amazon, and Uber, where engineers began to be measured on AI token usage and responded by inflating it. Meta had an internal leaderboard with status tiers such as "session immortal" and "token legend." Meta shut the leaderboard down in April, but according to Orosz AI usage still figured unofficially in performance evaluation. Engineers understood that a low token count was a bad signal, so they used AI for everything: rather than writing something by hand or reading documentation themselves, they asked the AI, partly just to burn tokens. AI is free for employees inside Meta, and people wanted higher bonuses. The code behind the incident was AI-generated, and AI reviewed it, in some cases "triple" reviewed it.

The layoffs made this worse. Meta laid off 10% of staff, about 8,000 people, on May 20, but the layoffs had been reported about a month in advance. Orosz said that during that month, worried employees increased their AI usage so their token numbers would not drop and mark them as layoff candidates. Instead of focusing on their work, including trust and safety work, people were focused on inflating their numbers.

The "psychosis" part concerned reassignment. Orosz said Instagram's trust and safety organization, built over seven or eight years and mostly based in London, lost about 40% of its people before May 20. They were told on a Thursday that starting Monday they would move to Alexandr Wang's AI organization to do manual AI data labeling: reviewing GitHub pull requests, adding tests, writing feedback, then moving to the next task. They were not given a choice. Orosz stressed that Meta had always treated engineers "like royalty" and let them choose their teams. They reported that close to 5,000 Meta developers now do manual labeling, and that employees joke this operation is bigger than OpenAI. After the layoffs and reassignments, most teams are less than half their former size, and some services no longer have anyone on call, which Orosz said had never happened at Meta before.

Orosz called this fully self-inflicted, coming from Mark Zuckerberg and Alexandr Wang, and said the message was that building a frontier model mattered so much that Meta would risk its business and its security. They described morale as worse than during the 2022–2023 layoffs, and added that in the US Meta records employees' screens and keystrokes to train AI. Engineers told Orosz they feel like discarded tools. Even people given large retention bonuses are interviewing elsewhere, not out of fear of being laid off but because they don't know whether they will be reassigned to data labeling, work they did not sign up for as professionals. Orosz said this might be the end of the Facebook engineering culture they had come to admire, built over 22 years. They acknowledged that Meta is an extreme case, but said it is happening right now.

Why "Everything Changed in the Last Six Months"

After the Meta story, Orosz laid out the plan for the rest of the talk: what has changed, a tour of what tech companies are doing, industry-wide trends, and advice for engineers and leaders.

To support the claim that coding changed dramatically in the past six months, they cited well-known skeptics. David Heinemeier Hansson (DHH), creator of Ruby on Rails, told Lex Fridman in October 2025 that AI wrote none of his code, partly because the models were not good enough. By February, when he appeared on Orosz's podcast, that had flipped: most of his code is now AI-written. Orosz emphasized that DHH cares deeply about software craftsmanship, is not paid by any lab, and told them that the models now write better code than he does. Simon Willison, whom Orosz described as one of the most quoted people on Hacker News and an independent engineer who experiments with every model, said that models released in November 2025, which Orosz named as Opus 4.6 and GPT 5.4, made agents genuinely useful. Given six months to absorb that shift, Orosz said, it is unsurprising that companies are now spending heavily.

Orosz then shared data from partners. Linear compared teams using AI agents to teams that don't and found that teams with agents ship five times as much code. Orosz noted that a similar increase previously took perhaps 20 years and this one took under two, while promising to come back to quality. Cursor reported that lines of code per developer per year rose about 2.5x, from roughly 3,000–4,000 to over 8,000, and that PR size tripled. Orosz multiplied these into roughly six times as much code, and acknowledged the audience's smiles: that probably means six times as many bugs. Cursor also reported a sharp rise in developers accepting AI changes without manual review, especially from January, when Orosz said Opus 4.7 and GPT 5.5 came out and companies started trusting Claude Code, Cursor, and Codex. Meta merging AI-reviewed code to production, they said, belongs to that same shift.

Anthropic and OpenAI: The Most "AI-Pilled" Companies

At Anthropic, Orosz spoke with Boris Cherny, creator of Claude Code, in February. According to Orosz, Cherny runs five parallel agents on a laptop at all times and ships 20 to 30 pull requests a day while leading Claude Code hands-on. Cherny told Orosz that PRDs, written planning documents, are dead at Anthropic and have been replaced by prototypes. Claude Code is 100% generated by Claude Code, and Anthropic overall is about 70–90% AI-written, with no target. People simply use it. Anthropic built Claude Cowork in ten days and it became a major commercial success. Orosz said sources at Microsoft told them Microsoft tried to build its own equivalent, and that Satya Nadella gave the team a month, but two and a half months later there is still nothing. Orosz used this to show that AI accelerates some companies while others, like Microsoft, stay held back despite having it.

At OpenAI, based on a conversation with the Codex team at the Pragmatic Summit in San Francisco, Orosz described an internal ChatGPT build with a "fix it" button: take a screenshot, say "fix this bug," and Codex generates a pull request that an engineer, or with safety nets even a non-engineer, can merge. AI code review happens in multiple tiers. Some code can reach production with AI review only, while critical-path code requires human engineers. Most developers run several agents at once. Orosz relayed the joke that engineers walk around with laptops slightly open so local agents keep running, and said that when they asked an OpenAI interviewee whether they had an agent running during the interview, the answer was "I had five." People attend meetings while their agents work.

Orosz noted that these are the most AI-committed people in the industry, who talk about AGI as a matter of when, not if. Most OpenAI staff barely write code by hand anymore. That changed around October. Some developers still wrote about 30% by hand then, and Orosz thinks that share is fading. The Codex team writes everything with Codex and told Orosz that taste, knowing what to build, is becoming very important. Codex also works on itself: because the team is mostly in one time zone, they kick off overnight runs where Codex tests itself and proposes improvements, which engineers accept or reject in the morning. In meetings and debugging sessions, they send voice notes to Codex at the start and have results by the middle or end. Orosz said it sounds like science fiction.

Cursor, Google, and Meta's Internal Tooling

Orosz visited Cursor's San Francisco office in October. They said Cursor went all in on agents in January, moving on from the tab-and-editor experience while still keeping it available. Cursor built its own coding model, Composer, which Orosz described as one of the few good coding models outside OpenAI and Anthropic, and notably cheap, which matters for costs discussed later. Orosz noted they have no affiliation with Cursor. Cursor runs tens of thousands of NVIDIA GPUs leased from Azure, AWS, and others. Inference used to be its biggest cost, but it now also trains models, turning it into a mini AI lab. Orosz mentioned that SpaceX seemed about to acquire Cursor. Everyone at Cursor is technical: Lee Robinson, who works in developer relations rather than as a software engineer, used Cursor to migrate all of Cursor's sites to a different CMS and published the cost.

At Google, everything is custom. The internal IDE, Cider, was web-based and is now a VS Code fork. Jetski is the internal version of Antigravity, integrated with the Piper monorepo. Critique handles code review, Code Search is best-in-class (Orosz said Sourcegraph was inspired by it), Borg is Google's Kubernetes, and Monarch its Datadog. Gemini is integrated throughout, so the internal experience is good. The problem, according to Orosz, is that Gemini is not as good as Opus or GPT 5.5. Engineers use Claude Code where they are allowed to, but only in limited contexts, so Google's AI adoption lags other companies. Orosz said the CEO has acknowledged this and Google is working on a better model.

At Meta, everything revolves around building a state-of-the-art model. Meta has an internal coding tool, Metamate, and a feature called "trajectories": GitHub-style commits show the exact prompts used. It was rolled out in December without warning people that prompts would be visible, so everyone could see, for example, staff engineers asking the AI to write a for loop. One developer told Orosz they began prompting in Polish so fewer colleagues could read it. Orosz believes Zuckerberg wants a model better than Opus 4.8, and that either Meta gets it within a few months or much of the company ends up deeply demotivated.

Uber: What Serious In-House AI Tooling Looks Like

Orosz, a former Uber engineer, used Uber to show how much in-house AI tooling a large company builds. Uber has about 3,000 engineers and about 20,000 other employees. A developer experience team of around 20 people, now effectively an AI experience team, has built:

  • An internal MCP gateway for discovering and registering MCP servers.
  • An Uber Agent Builder and Agent Studio, a no-code, drag-and-drop way for the 20,000 non-engineers to build agents, plus a registry for them. Orosz compared it to OpenAI's public Agent Builder.
  • An AI CLI, which Orosz called "the Claude Code for Uber," integrated with internal systems and multiple models.
  • Uber Minion, background agents at scale, similar to Cursor's background agents but integrated with Uber's monorepo and experimentation system. Orosz said developers who can use Claude Code often prefer Minion because it works better and faster. For example, it analyzes a prompt and suggests how to rewrite it for faster, cheaper, better results, which Claude Code doesn't do yet.
  • A code inbox that surfaces which of the flood of mostly AI-reviewed changes actually need a person's attention, with smart assignment and SLAs that escalate to someone else if a reviewer doesn't respond within about a day, like on-call tooling.
  • Risk profiles that flag changes that look risky.
  • uReview, an internal AI code reviewer comparable to CodeRabbit or Sonar.

Orosz said other large companies are doing the same, each with a dedicated infrastructure org: Stripe (Minions, Toolshed, Blueprints, devboxes), Ramp (Inspect, Glass, Dojo, Sensei), Shopify (Sidekick, an LLM proxy, a dev MCP server), Airbnb, and others. If you're proud of wiring an agent into Slack, they joked, that's cool, but this is the next level.

Startups and Traditional Companies

Among startups, Orosz sees the familiar pattern: agents write and review code, and there's plenty of Slack-based creativity. At one startup that had raised a $70 million Series B, someone jokingly told an agent in Slack to "fix all bugs in the codebase." It came back having found four critical authentication issues that left a back door wide open. Orosz also sees startup engineers vibe-coding their own replacements for SaaS tools, which they read as engineers having fun rather than a real business trend, and which they don't see at large companies.

The surprise, for Orosz, was traditional companies. They lack Uber-style platform teams but are not really lagging. Cisco rolled out Codex to 18,000 engineers in January, when Codex was still small, and uses it for complex migrations. JPMorgan Chase built a multi-agent framework, specialized agents labeling customer interaction data with evals and judge-based aggregation.

Trend: Individual Speed-Ups Don't Add Up to Team Results

The first industry-wide trend came from Laura Tacho, formerly CTO at DX and now leading developer experience at AWS, whom Orosz messaged the night before the talk. Tacho told them that many organizations see individuals doing great but no improvement at the team level, because they treat AI as a personal productivity tool, which she calls "individual speed-up juice": email summaries, Slack automations, even code generation. The companies seeing team-level results start from a business outcome instead, such as deploying faster, shipping more features at the same quality, or improving quality.

Orosz offered Spotify as an example, based on lunch with its CTO about six weeks earlier. Spotify's bar for AI use is that quality must stay the same. As a result, it isn't seeing huge output increases, but it has built internal quality-checking tools and deliberately rolls AI out more slowly, a contrast with Meta. Tacho's framing is to build agentic systems that reduce handoffs, make information easier to find, and remove friction while maintaining quality, which Orosz said few companies do. Tacho maps companies on two axes: AI usage (individual vs. team) and decision-making (simple automation vs. agentic systems). Most companies sit at individual/simple automation, and most want to reach team/agentic systems. Getting there, Orosz argued, requires what Uber did: building integrated systems with your own engineers over time and at great cost. You can't simply buy Claude Code or Cursor and get it, whatever vendors say.

Trends: Token Maxing, Tool Addiction, and Vanishing Middle Management

Orosz said token maxing, burning tokens on purpose without shipping value, is going out of style but persists under pressure to look productive, especially at US tech companies that "don't really care about budget until they do." They said it has happened at Meta, Amazon, and Microsoft, which they said still runs an internal leaderboard. They also described the pricing of AI tools as somewhat addictive: start at a $10 or $20 plan, hit a generous limit, upgrade to $100 or $200, then feel pressure to use up the allowance, and eventually end up on API pricing. For some people, especially in the first months, prompting feels like gambling ("one more prompt"), with people sleeping poorly and waking up thinking about their agents.

Another trend is the flattening of middle management: managers laid off, reassigned to individual contributor work (as at Meta), or told to be hands-on and manage less, often with AI given as the reason. Orosz pushed back on the popular disdain for middle managers and directors. In their experience, good middle managers are very technical and could be hands-on but choose not to. They watch what's happening and intervene, for example noticing a run of outages and pulling people into a task force to fix the underlying system, where engineers alone might just pile on. Good engineering management improves engineering culture, and Orosz said that removing it will make culture decline, which they called "a fact as far as I'm concerned."

At the same time, CEOs and CTOs are returning to coding. Guillermo Rauch, founder and CEO of Vercel, told Orosz he sees executives coding again "with a fury," including public-company CEOs messaging him excitedly about using Vercel or Claude Code. Orosz pointed out the combination: fewer middle managers to protect engineers, and executives vibe-coding things they think are complete when they aren't.

Trend: The AI Bill Arrives

Orosz called rising AI costs a "mega trend" they noticed only a week or two before the talk and had written about for subscribers the previous week. On the morning of the talk, Sam Altman posted that AI budgets had become a big issue for some companies. Orosz said they asked contacts at OpenAI whether Altman reads their newsletter, and heard that someone had posted it in Slack and he had read it. They also mentioned a joke circulating on Reddit about a partner discovering $15,000 gone from a shared account.

According to Orosz, Anthropic has moved enterprise customers onto API pricing, so anyone who isn't a startup or individual no longer gets discounts. GitHub Copilot made a similar change on June 1, two days earlier, and users are furious that budgets, $200 or whatever they were, that used to last a month are gone in three days. Uber's CTO said in March that the company had burned through its whole annual AI budget. Uber now caps AI spending at $1,500 per engineer per month, after which engineers use free models. Orosz said their research shows many companies doing something similar, some capping as low as $200 before falling back to zero-cost models on Copilot. Costs that approach an engineer's salary are something no one wants to pay, "no matter what the AI labs say."

Craft Trend: Quality Is Dropping Everywhere

Turning to the craft itself, Orosz said quality has dropped across the board, starting with their own experience. For about a month, as a paying user, every time they opened claude.ai and started typing, a React lifecycle refresh wiped out what they'd typed. They screen-recorded it and tweeted about it. A product manager replied that it was great feedback and would be fixed. Orosz read that as an admission that Anthropic had no idea millions of people hit the bug daily and wasn't dogfooding its own product. They said a bank would do better, even though Anthropic did eventually fix it.

OpenAI bragged about building its Agent Builder in six weeks with one engineer using Codex, but Orosz said users hit many problems at launch and its forum filled with unresolved complaints. Three months later, one user wrote that they had been bullish but P0-type bugs weren't getting fixed and it felt like abandonware. AI made it faster to build, Orosz said, but not better. At AWS, an engineer let the internal Kiro coding tool make changes, and the agent chose to delete and recreate an environment, causing a major outage. AI-generated code also took down part of Amazon's flagship retail site. Orosz framed both cases as over-reliance on AI or not caring about quality, and said Amazon now requires a senior engineer to review any AI-generated change, because junior engineers would just approve it.

Orosz highlighted Dax Raad, founder of OpenCode, an open-source AI harness with nearly a million daily active users that has grown about 10x in four or five months. On Orosz's podcast the week before, Dax said the team was shipping more hacks where they should have redesigned systems from scratch, and that "our judgment is off." Dax also observed that no competitor is beating OpenCode by using AI better, admitted "I don't think we're using AI that well," and said the team is telling itself to use less AI, to think more, and to build fewer things that matter. Orosz noted the irony of an AI tool founder saying this, and suggested OpenCode remains one of the highest-quality harnesses precisely because it slows down.

Craft Trend: Everything Is Broken, and Reviewers Are Burning Out

Orosz's example of things breaking was GitHub, where two weeks earlier pull requests disappeared for 8 to 12 hours. An unofficial uptime tracker that aggregates reported outages estimates that GitHub doesn't reach even one nine, meaning some part of it is down around 10% of the time, which Orosz called "absolutely unserious." GitHub's leadership shared unpublished numbers with Orosz, attributing the problems to a 3x load increase over two years. Orosz said they don't buy it. They did not claim AI-generated code was to blame, but argued that if GitHub can't handle 3x growth in two years, with more to come, something is wrong, since startups absorb that kind of growth easily. Mario Zechner, creator of pi, told Orosz that software has become a brittle mess, with 98% uptime feeling normal and UIs full of strange bugs. Zechner conceded this predates agents but said it feels like it's accelerating. Orosz added that they recently had a serious software problem with Magyar Telekom, though they doubted AI was involved there.

The last craft trend is that AI "slop" is burying the engineers who still care. There are far more pull requests, mostly AI-generated, and most developers see that the AI reviewer passed a change and give it a thumbs-up without reading it. The minority who still review properly, catching bugs, pushing back, and spotting duplicated code, are overwhelmed and burning out. At performance review time they aren't rewarded, because they aren't seen as the people shipping features. Some quit. Dax told Orosz that OpenCode is hiring many of these people, who left because they were the last ones keeping things alive and no one cared. With engineering management gone or less hands-on, there's no one left to care.

Kent Beck summarized it for Orosz: "We're accumulating code faster than we accumulate trust." Code has to be understood and trusted, and there's no time for that right now. Beck also said AI amplifies experience, so seniors gain the most and judgment is rewarded. Orosz cited Hillel Wayne: the only people successfully generating working TLA+ specifications with AI are TLA+ experts who spelled out exactly what to generate. The same applies more broadly, Orosz said. A junior engineer who has never built an iOS app can prompt one into existence, but it won't be maintainable. Old patterns are returning as a result. Dax told Orosz that OpenCode relies on domain-driven design and verbose guardrails, because agents are like new junior engineers who need guardrails. Enterprise patterns that fell out of favor for being long-winded keep agents in check, and Orosz said, "dead serious," that it may be time to dust off the design patterns books.

Advice: Ship Only What You Can Verify, and Keep Learning

Orosz's first piece of advice is the talk's title. Cap your daily agent usage at what you can review or verify. Verification doesn't have to mean reading every line. Peter Steinberger, creator of OpenClaw, told Orosz he ships code he doesn't read, but builds his own verification systems, thinks at the architecture and module level, and has the AI draw diagrams for him. Either way, don't ship more than you can verify.

Second, tech debt is now cheap to remove: have agents kick off agents to clean it up, and become your team's "chief tech debt remover." Orosz said if you aren't doing this, you aren't using AI efficiently. Third, experiment, since there's no one-size-fits-all. Share approaches with friends. Mitchell Hashimoto, creator of Ghostty and founder of HashiCorp, told Orosz on the podcast in March that he keeps exactly one agent in the background: when he's coding, the agent plans, and when the agent codes, he reviews. He doesn't use multiple agents. Orosz contrasted this with people running five agents, which they personally don't understand how anyone manages, and stressed that Hashimoto cares about high-quality software, not AI for its own sake.

Fourth, spend more time thinking and understanding. Orosz relayed Dax's quip that he used to spend 95% of his time thinking and 5% coding, and with AI now spends 96% thinking and 4% coding. Orosz said most people won't have that luxury but it's a useful frame. Finally, citing Addy Osmani: don't outsource learning. It's too easy to let AI fix a bug while your mental model stays unchanged, trading long-term capacity for present-day speed, and the tools won't stop you. Whenever you use an agent, learn something.

The Job Market and Future-Proofing Your Career

On the job market, Orosz said the global picture looks okay, based on a Pragmatic Engineer deep dive using Indeed data, which they consider reliable. Top tech companies are hiring more than before, though conditions aren't as good as they once were. Orosz cited a 20% change in software engineering postings in the US and UK (introduced as bad news but described as an increase, so the framing in the talk is unclear), decreases of 13% in Germany and 10% in France compared with two years ago, and flat numbers in Canada. They had no Hungarian data but suggested Hungary's close ties to Germany may mean a similar trend. At top US tech companies, AI engineering hiring is booming and now makes up about 10% of all software engineering roles. Orosz defined it as building RAG, evals, and systems that use LLMs.

For future-proofing, Orosz recommended building things on top of AI and LLMs, whether a side project like a podcast recommender or an internal tool at work, and suggested Chip Huyen's AI Engineering as a practical book. They urged engineers not to outsource their thinking, to become product-minded (pointing to the book and their blog post The Product-Minded Engineer), to talk to product managers, and to understand the business. They also recommended becoming a domain expert. At an agriculture company, understand agriculture, because many engineers exist but few have talked to farmers. At an automotive company, talk to mechanical engineers. That expertise makes you valuable during downsizing or a job change.

For engineering leaders, the advice is to be or stay hands-on, "yes, again, otherwise you will be out this time." Orosz said AI makes this easier, since it can explain code and help you contribute. They had heard of companies where five of the top hundred committers are product managers or engineering leaders. Leaders can also help integrate AI at the systems level to remove friction. But they should expect to do less people management, because the business expects it, and those who love people management will either burn out or do less of it. For engineers, that means less career support and probably smaller raises for a while.

Closing: Change This Fast Is Overwhelming, and That's Okay

Orosz closed by noting that Martin Fowler, Grady Booch, and others told them the industry hasn't changed this fast since the 1960s: in about 12 months, AI went mainstream across coding tools. Orosz said feeling overwhelmed is fine and that they sometimes still are, given how fast and unpredictable the change is. Their final suggestion was to pause occasionally, acknowledge that you are keeping up, and ask how to make your work more sustainable, how to produce higher quality, and what you can automate. Then make a change, and repeat.