Brian Scanlan on Running a 15-Year-Old Rails Monolith as an Agent-First Engineering Organization
Ruby on RailsIntercom runs a Rails monolith that is about 15 years old and contains millions of lines of code. It also describes itself as "technically conservative." In this episode of On Rails, host Robby Russell asks Brian Scanlan, a senior principal engineer on Intercom's platform group, how those two facts fit alongside an aggressive, company-wide bet on Claude Code. Scanlan's answer is that a boring, well-understood foundation makes the bet easier. The conversation covers what broke as a result, including CI, code review, and the limits of the models, and how Intercom has been rebuilding its development lifecycle around agents. The conversation was recorded in mid-April 2026.
Why Intercom stays on Rails: being "technically conservative"
Scanlan jokes that the monolith is simply too big to break into microservices, "obviously something we would love to do." The serious answer is the usual one. Rails is expressive, easy to get started with, and light on boilerplate. That lets developers focus on building product instead of learning huge frameworks or building complex microservices.
Scanlan ties this to a broader engineering principle. Intercom calls itself a principles-driven organization, and the principle that has served it best is being technically conservative. In practice this means large applications into which the platform teams put "all of the smart stuff." Product engineers then spend as little thought as possible on "gunge and undifferentiated heavy lifting" and more time shipping product.
Asked how this relates to "boring," Scanlan cites Dan McKinley's Choose Boring Technology. They re-read it periodically and call it "gold." In their summary of McKinley's argument, if you build on well-understood components, scaling work gets absorbed into the shared effort of keeping, say, the MySQL layer or the Rails app healthy. If a team picks a bespoke technology that fits its problem perfectly, it now owns a pet that it has to feed and maintain. Scanlan concedes that approach is better than doing nothing, but holds that in most businesses, investing deeply in a small set of technologies gives everyone more throughput and higher quality. "It's hard enough to scale one database," Scanlan says. Scaling three because different teams picked different ones would be far worse.
The platform group and how it partners with product teams
Scanlan has been at Intercom for 12 years. The platform group is responsible for availability, performance, cost, security, and developer productivity. That covers the monoliths and developer environments, and Scanlan does everything from on-call to interviews to architecture. Scanlan says the job lately feels like being "some sort of Claude Code evangelist or whisperer." The group has also been especially influential in Intercom's AI rollout.
The platform group owns Rails upgrades, but Scanlan draws the line with product teams much deeper than the framework version. The group owns the interfaces product engineers build on: async workers, web servers, the "glue" patterns, observability, and everything from the AWS accounts up to large amounts of Rails code. It pairs with product teams on operational issues, database upgrades, and query optimization. It does not treat the Rails app as a black box that people bring problems to. The group owns the problems, and will sometimes change how features work to keep them performant, cost-effective, or upgradable.
Robby asks whether there is a DBA-style gatekeeper for schema changes. Scanlan says no. The goal is for product engineers to use ordinary Rails migrations and mechanisms, and have them work in production as easily as they would in a simple environment. The platform teams are staffed mostly with former product engineers, not sysadmins or DBAs, which Scanlan finds ironic given their own SRE background. People typically drift into a niche such as database performance or deployments and then ask to join. Intercom generally doesn't hire for specialist knowledge, or even for "10 years of Rails." It prefers generalists who are willing to go deep, while keeping enough experienced people on the bench.
What keeps Scanlan up at night: not moving fast enough on AI
Scanlan's current worry is that Intercom isn't adopting AI fast enough. Having talked to many companies, Scanlan guesses Intercom is in "the very high percentiles," but says that isn't good enough. The anxiety is about headwinds. A greenfield competitor doesn't carry 15 years of complex features and many different code styles. Scanlan considers it unacceptable for Intercom to be held back because a million lines of code means Claude Code takes five times as long to get something good.
Scanlan also believes that even if the tools stopped improving, Intercom is on a path to assembling, "Voltron style," enough Claude Code skills and guidance to handle the vast majority of what a product engineer does. The impatience is about finishing that work while keeping a very high quality bar. Scanlan isn't very worried about going down wrong paths, since you can notice when things go badly. The aim is to get it right the first time, and changing how people work is a huge amount of effort. "Turns out that humans do a lot of work," Scanlan says.
When CI blew up
Robby had read Scanlan's tweet thread on how Intercom uses Claude Code, and was struck by how much pressure the CI system had come under. Scanlan frames it as a problem of success. They like to say that when a scaling problem hits, the first step is to throw a party, and the second is to figure out the fix.
The success was a rapidly rising rate of pull requests. The monolith's test suite has hundreds of thousands of tests and would take about two and a half days of wall-clock time to run serially. Intercom parallelizes heavily, makes hundreds of changes a day, and ships to production up to 100 times a day. The company is so dependent on this that being unable to ship is treated as a P1 event, whether the cause is a flaky test or a stuck pipeline.
Two things went wrong. The first was flakiness. About a year earlier, Intercom had adopted Shopify's Rotoscope to identify changes where only a subset of tests needs to run. Scanlan says this saves a great deal of compute and cost, but little time, since the slowest test and the setup and teardown still dictate duration. With higher velocity and Rotoscope, retries and flaky tests grew. Intercom automatically retries failed tests. A genuinely flaky test normally fails under one seed or ordering and passes on the next run, which is tolerable when you ship all day. Past a certain critical mass, though, the flaky tests started "tripping over each other." The second problem was cost, which Scanlan says was growing faster than the business, and the business itself was growing fast.
There was no single fix. The team resized hosts and put everything in scope, including temporarily stopping its aggressive use of AWS spot instances. Scanlan explains that spot hosts can be reclaimed at any moment in exchange for a lower price, which suits tests because you can simply rerun them. They still need tuning, for example so that a host signaling shutdown stops picking up new tests. Pausing spot usage stabilized things for a while.
The rest was "the classic boring scientific method": use observability to find the bottlenecks. Using Honeycomb tracing, the team found that most of the time went into loading factories. Heavyweight "semi-god objects" such as conversations and workspaces were set up hundreds of thousands of times across the suite. A deep conversation factory that worked for almost any test could be replaced with a lightweight version that only does the complex setup when needed. Scanlan says dozens of fixes like this, aimed at the biggest sources of time and flakiness, got CI into a better state. None of it was rocket science.
The mandate: double engineering throughput
Scanlan says throughput started rising around the previous summer. That coincided with CTO Darragh publishing a goal to double engineering throughput, measured crudely as pull requests per head in R&D. Scanlan acknowledges that any metric stops being a good measure once it becomes a target. Still, if AI tooling makes producing code much easier, a natural rise in PR throughput is a reasonable expectation. The team already knew CI had low-hanging fruit, so its trouble wasn't a surprise.
Intercom also invested in linting. It already had a decent RuboCop setup and GitHub Actions checks, but expanded them on the expectation that more code needed to be better code. Claude Code "still can get stuff wrong," Scanlan notes. Without guardrails it can produce confusing code or code "inspired by code we wrote 15 years ago," which is worse. The number and sophistication of Intercom's custom cops are now much higher. Scanlan credits Claude Code with being good at writing cops, whose syntax can be intimidating.
One effort had mixed results. Scanlan once counted about 10 HTTP client wrappers in the monolith, including Typhoeus and various per-service wrappers. The team started consolidating them, on the reasoning that what's good for humans is good for LLMs and agents should know the one way to do things. The work petered out because its impact was hard to prove. Scanlan concluded it's easier to tell Claude Code which HTTP client to use, or "do not use metaprogramming," than to perform surgery across the whole app. Many old patterns stay in place, and the guidance keeps the agent from copying them.
One tool, agent-first work
Robby asks what the workflow looks like when an engineer picks up a ticket. Scanlan starts with the mandate: Claude Code is Intercom's single tool. Without one platform, it's hard to build up the sophistication of repeatable skills, whether for Rails upgrades, flaky tests, outages, or new features. Scanlan cites Greg Brockman's post saying all technical work at OpenAI would be agent-first by March 31st, and says Intercom has the same principle, with OpenAI "a little bit ahead." If you aren't opening an agent first, for outage response, planning, research, or writing code, "you're kind of doing it wrong."
By Scanlan's figures, well over 90% of code changed daily in the Rails codebase is generated by Claude Code, and some days 95% or higher. The vision is for people to stop reading code and close their IDEs. You tell Claude the problem, it picks the skills, and it produces the right code the first time. Scanlan says they rarely open an editor anymore. Many people still use editors, invoke Claude from VS Code, or inspect the output, but Scanlan thinks most of Intercom is well along a maturity curve that runs from tab completion, to the agent writing most of the code, to no longer looking at code at all. Scanlan doesn't write giant specs. Their preference is interactive, conversational, interview-style work in which the plan develops as it goes.
The exception is "artisan" work. Scanlan's metaphor is that the software factory, the assembled Claude Code skills, produces IKEA furniture. Ornate, novel, or unusually high-quality work that the factory hasn't seen still calls for skilled humans. UI is an example: Claude Code still can't reliably notice that something is "off by two pixels," so visual work benefits from a human and an editor in the loop. Scanlan describes their own job right now as filling the gaps. Whenever the code isn't right the first time, they figure out why and fix it.
Local agents today, remote agents next
Asked whether autonomous agents pull tickets off a backlog, Scanlan says the work is mostly driven from local Claude Code. When paged into an incident, Scanlan habitually opens Claude and tells it about the incident, and muses about automating that step. Intercom is building its own remote agent system, which effectively mimics its local Claude Code setup. Scanlan mentions published work from Stripe and Ramp and various startups as other options. Remote agents are needed for multiplayer, long-running, large, scheduled, or automation-triggered work, such as joining incidents or Slack channels on their own. Scanlan notes, though, that recent features like scheduled tasks have narrowed the advantage, and that "just local Claude Code is a perfectly good way to get through a lot of work."
Code review is the new bottleneck
Robby describes clients whose developers now spend more time reviewing the growing number of PRs, and who wonder whether humans should review at all. Scanlan agrees that code review is the current bottleneck, and says that's exactly what increasing throughput is for: you find the next bottleneck and fix it. Intercom's goal is for the majority of pull requests to be approved automatically with no human approver. Today about 20% of Rails app PRs are fully auto-approved by a Claude Code–based system that identifies "safe PRs." The target is over 50%.
Scanlan traces this back several years to Dependabot. Intercom used to open issues for teams to handle Dependabot security PRs, and the teams ignored them, so hundreds or thousands piled up across a long tail of applications. Scanlan questions what value a human adds when upgrading, say, the JSON gem in a huge codebase. The human isn't really re-verifying anything and is relying on unit, integration, and smoke tests and post-deploy exception monitoring anyway. So Intercom stopped putting humans in the loop and began auto-merging Dependabot PRs in batches on Tuesdays, Wednesdays, and Thursdays, visible in a Slack channel and never at 2 a.m. A handful of "load-bearing" dependencies are excluded based on experience: the AWS gem and Rails itself, and LangChain in Python. Scanlan says outsiders are often aghast at this trust in CI. Their view is that the value of a human in that loop is very low, and people shouldn't be asked to do low-impact work just because it's the process.
Intercom also has deterministic, path-based auto-approvals. Changes that only touch specs, documentation, and similar files ship without human approval. Scanlan says all of this is written into Intercom's SOC 2 and ISO 27001 compliance. "Despite rumors to the contrary," these regimes don't require humans to approve each other's work. They require documented processes, controls, risk mitigation, and audit trails, all achievable with automated review. Scanlan sees no reason a mix of agents and rules can't produce faster, more repeatable, and better review processes.
"Putting manners on the bots"
The first attempt at AI code review, simply asking Claude Code to review PRs, produced reviews Scanlan calls terrible even though they looked good. LLMs are good at producing things that look like good work, are eager to say anything, and can make a big fuss over minor points. The fix was to "put manners on these bots." The reviewer was told not to get involved unless something is very impactful, to keep it short, and to say nothing if it has nothing to say. It was also given examples of the kinds of issues Intercom wants caught.
Because Intercom has a huge history of changes, it could generate many candidate reviews and compare them with reviews by its best Rails engineers. Those engineers then graded the LLM output, creating a feedback loop instead of relying on one-shot reviews. Scanlan says that once the noise dropped, people started acting on the feedback. Noisy reviewers get ignored.
Approvals followed the same process: backtest historical reviews and approvals with your best people, and decide which criteria you're comfortable with. Intercom combined historical data, expert input, and common-sense rules. A good candidate change isn't hundreds of lines long, sits behind a feature flag where possible, and avoids certain code paths. The agent's job is to answer "is this safe to approve?" Scanlan acknowledges that reviews also serve as a teaching aid and catch other problems. Intercom's bet is that it can push more throughput through the system, learn from what goes wrong, and improve its checks, its codebase, or even its deployment system, which Scanlan calls very simple and possibly in need of more features.
Why keep humans involved at all?
Robby pushes on an apparent contradiction. If one engineer's Claude writes the PR and another engineer's Claude reviews it, and nobody is meant to read the code, why not have a single engineer chain skills for submit, fresh-context review, fix, and merge? And how does accountability work if engineers are told not to look at the code?
Scanlan calls it tough. Code review has been the default gatekeeping mechanism and a way to see what's changing, and some things in the code and infrastructure Scanlan still wants to know about. But much of that visibility can be handled reactively with agents. A team working on the Intercom Messenger, for example, can ask Claude for a weekly summary of what changed in its area. Blocking people from shipping is "a pretty extreme solution" for that need. Where specific concerns exist, such as Intercom's hottest code path, the review agent can refuse and demand a human. Scanlan calls that the "ejector seat." Education and awareness can also happen after the fact.
Scanlan's core argument is that most code "is just not that interesting," especially once it has passed heavier linting and tests and been pushed by the auto-approval criteria into small, tested changes. What would have been risky before is, in their view, far less risky now. Scanlan also describes a deployment agent in development that would monitor 100% of deploys. It would read the code and determine which metrics to watch, the way you'd have your best engineer watch a Rails upgrade "like a hawk," but for every change. Scanlan presents this as a capability Intercom never had before, and one that offsets the risk of moving more changes with fewer humans in review.
The skills architecture
Scanlan says product engineering "might actually be the least interesting part" of the story, but first describes the plugin structure, which is shared between R&D and the rest of the company.
A base plugin provides telemetry and safety. Claude Code hooks send metadata about every session and skill call to Honeycomb, so anyone developing a skill can explore how it's being used. Full session transcripts, usually large JSON files, are copied to S3 and lightly anonymized. Scanlan admits there's only so much anonymization you can do. Sessions from groups such as legal and finance are excluded. The purpose is quality control and a feedback loop, since "there's gold in there" about how people work and whether things are working.
Core plugins cover general developer work used across Intercom, such as fixing a flaky spec or opening a pull request. The bar is very high. Core skills must have evals, must follow both Anthropic's skill guide and Intercom's internal guides, and are improved with "skill improvement skills." Skills graduate into this tier, and once there, the team commits to maintaining them. Core skills should be unopinionated and usable by anyone.
Team plugins have "exploded." They fit specific workflows and can carry a lower quality bar and more opinionated setups, because overfitting to a team's preferences is the point. Some are high quality but deliberately restricted. The Rails upgrade skill is one: Intercom has been running bleeding-edge Rails in production with weekly upgrades, and has encoded the repetitive steps into a continuously improved skill. Only one team should be upgrading Rails, so the skill is hidden from everyone else.
Below that is a long tail of personal skills and marketplaces, and Intercom is deliberately liberal about letting people build and distribute them, partly so people learn. Scanlan favors small, discrete, testable skills over monolithic ones, which can be composed from smaller skills. Intercom accepts overlap. Ten different "investigate a bug" skills can coexist, and the team then uses session data to find which parts actually get results and pulls out the best. Momentum matters more than going slowly and hoping the good stuff emerges on its own.
Skills, rules, guidance, and hooks as different levers
Scanlan describes skills, rules, guidance, and hooks as levers suited to different goals, and illustrates with one of Intercom's earliest skills, for pull request descriptions. Out of the box, Claude Code "regurgitates" what the code does, which Scanlan calls useless. PR description quality fell as adoption grew, because the important part is why the change exists. Even a link to a GitHub issue is more useful than an English paraphrase of the diff.
Intercom first added a rule that calls a skill. The skill reads the session data to establish the purpose of the change and asks the user if it can't tell. That helped, but the rule wasn't being invoked consistently, so the team moved to a hook to force the behavior. The general method is to experiment, use telemetry to see whether behavior changes, and ask people why a skill wasn't used: are they on the right Claude Code version, which plugin versions do they have? Scanlan calls it a "huge tech support kind of thing." Intercom is at the mercy of whatever Anthropic ships next, but Scanlan considers the building blocks good enough.
A thousand weekly users, most of the value outside engineering
Intercom has over 1,000 weekly Claude Code users, and adoption went "viral" outside R&D once the tool was connected to Snowflake with good support. Sales, marketing, customer support, and other teams began doing work they had long wanted to do: answering their own questions and encoding their workflows as skills. Scanlan credits three things: picking one tool, providing access with good controls over Snowflake data, and a core team setting people up with MCP configs and lots of technical support. "If that's not there, then you're going to have a bad time." Scanlan reports people saying in internal channels that they're doing the best work of their careers and no longer wait months for an analyst to run a query.
Robby asks whether this lets Intercom avoid building internal admin tools and reports. Scanlan says writing code "might be the least important thing" happening with Claude Code at Intercom. The main non-R&D use is analytics: making sense of data in Snowflake, Tableau, and elsewhere, where people previously didn't know which fields to trust and "Snowflake doesn't actually know." With good guidance and access, people ask a question and get the answer ad hoc. Scanlan suggests that many dashboards, and the teams maintaining them, are "almost obsolete in a certain way," since their real job was answering questions. Scanlan considers this a much bigger deal than reliably producing database migrations.
The Rails console MCP, and why non-engineers use it most
Most integrations are standalone local MCPs for individual services, which means authenticating "to a million different things." Intercom has also been consolidating functionality into the Rails app. Scanlan recalls that when they joined, customer support staff queried production directly in the Rails console. As Intercom grew and "became more boring," the console was locked down to engineers, for good reasons.
One of Intercom's philosophies is that anything you can do on your laptop, your agent should be able to do. One day Scanlan was running console queries and pasting the output into Claude, and realized they were acting as a proxy for the agent. Intercom was already building API versions of internal tools like customer lookup, so Scanlan built an MCP that lets Claude Code send arbitrary Ruby to the Rails console, layered on the console's existing safeguards. Scanlan was a little apprehensive and launched it quietly. The top five users, each running hundreds of queries of ad hoc exploration and often combining the console with Snowflake data, were all non-engineers. "I felt so proud," Scanlan says. Design managers and senior product directors were "unwittingly" running fairly complex Ruby because Claude chose the console as the best place to answer their questions. Scanlan suspects Intercom's head of product wouldn't even know they had used production access.
Technically, the MCP takes a string of Ruby as an API parameter, which can be a multi-part, lambda-like block. It runs checks against it, executes it in the production environment with the same functionality as SSHing into a host and opening a console, returns the output, and logs everything to an audit trail. The experience is worse than a real REPL. There's no tab completion, and each call may land on a different server, so the code has to be right the first time. Claude often sends invalid code, but when it gets it right, Scanlan says, "it's amazing to see." Claude learns what to write from a skill that accompanies the tool with examples. Users are also usually running Claude inside the Rails codebase, so it alternates between exploring code and running queries in production. The tool is currently read-only. Scanlan isn't opposed to writes on principle, but nobody has asked for them. If something needs to be written regularly, the answer is to build an API endpoint, "one Claude Code session," with proper validation.
On security, Robby asks about a hacked laptop exposing all production data. Scanlan answers that these users already had access to this data, and Claude acts as a proxy. From a strict compliance standpoint, what people can access doesn't change. The volume of access does grow, and laptops now hold more credentials and active sessions, so detection matters more. Scanlan also calls securing agents in general "very much an unsolved problem." Intercom's mitigations include governing which MCPs people can connect to and where plugins come from, detection on laptops, and distributing plugins and locked-down Claude configuration through internal IT systems. Scanlan admits this doesn't stop someone from downloading Codex and bypassing it all, but says it makes the default safer and avoids encouraging shadow IT. With recent news about new models and security threats, Scanlan expects things "to get more exciting for security teams," and says the answer is doing the fundamentals well.
Teaching the terminal, again
Asked whether non-technical staff are intimidated by a terminal UI, Scanlan recalls getting into computing about 27 or 28 years ago, running student-run Unix systems at university around 1998. They hosted newsgroups, IRC, instant messaging, and email: "the best social network" on campus. Thousands of non-technical students were "banging down our door" to learn the terminal because they saw friends doing cool things with it. Nearly 30 years later, Scanlan says the job is the same.
There's also bottom-up enthusiasm. Intercom's VP of operations, after teaching colleagues, used Claude Code to build a small guide for non-technical people getting started, warning that the terminal is "real nerd stuff" but manageable. Scanlan's conclusion is that motivated people get past the terminal quickly once they see productivity gains, even if few will go as far as installing Oh My Zsh. Robby, who created Oh My Zsh, recalls doing it to make designers and front-end developers comfortable with the terminal.
Where AI still fails: a Datadog false alarm
Robby asks about deterministic versus open-ended work. Scanlan says that even in Intercom's well-guided environment, Claude Code struggles with open-ended problems, and gives an example from three or four weeks earlier. One of Intercom's most important alarms fired on a weekend. It's an SLO-style alarm on inbox usage that fires when customers' reply behavior deviates from statistical predictions. Past causes have included slow databases, internet issues, and Carnival in Brazil.
The on-call engineer did what Scanlan would do and brought Claude in. Claude produced many plausible theories, including blaming changes from the previous day, even though the drop was short and had recovered, which didn't fit a latent bug. The theories sent the engineer down several paths, and they settled on one. A couple of hours later Scanlan dug in and found that the real cause was an eventual-consistency issue in Datadog's metrics, not an obvious conclusion. Scanlan admits to being pleased for a few minutes to still be able to "outsmart" the LLMs.
The lesson Scanlan draws is not about the individual. The engineer was experienced at Intercom and, in Scanlan's words, one of its best AI engineers. The lesson is that in open-ended situations without a clear goal, the model gives confident, plausible answers, and it takes real skill to say "the vibes are wrong." Novel cases can be encoded so Claude checks them next time, and Scanlan thinks that eventually covers most edge cases. But being led down red herrings is worse than Claude admitting it has no idea. Repeated at scale, it can build wrong beliefs about production or leave root causes unsolved. Scanlan acknowledges humans do this too. Asked whether incidents get a second look by process, Scanlan admits to "cheating": they join nearly every incident channel and apply "the eye of Sauron," which they concede may not scale at every company.
The airline pilot metaphor
Scanlan's model for the future is the commercial airline pilot. Pilots land mostly on autopilot and monitor and manage it. They also know they'll be tested in simulators without it, take pride in the craft, and deliberately disengage it for some share of landings to stay sharp. Scanlan is "extremely bullish" that Claude Code will eventually resolve most incidents faster than Scanlan can, and calls that the bar, while saying Intercom is still a good way from it. Even then, Scanlan plans to race it, or disable its ability to act, and compare.
The reason is that running the system is Scanlan's job, and a factory owner has to understand the output deeply, "take the IKEA furniture home and use it and assemble it." Even for well-understood, deterministic jobs like Rails upgrades and flaky test fixes, which Scanlan considers about 90% solved with Claude Code, people still need to turn the autopilot off at times. Without the expertise to judge the LLMs' output, "you lose control of the system."
Robby raises aspiring entrepreneurs who want to build apps entirely with fleets of role-playing agents. He also recalls learning Rails when it was considered risky and figuring things out as problems arose. His view is that the industry will work these questions out together as it encounters them.
The 2x goal was met; now for more
Scanlan reports that Intercom met its goal of doubling R&D throughput within about a year, roughly two weeks before the recording, and now intends to double it again. Scanlan now considers 2x "kind of underwhelming" given where models and harnesses are going. The organizational change of rewiring how hundreds of people work is years of work that Intercom is trying to "speed run." Scanlan compares the shift to compilers: as more work moves to agents, people will think less about low-level details and more about inputs and outputs, while still being able to go deep. Scanlan asks, "why can't it be 10x?" and thinks that's realistic.
Token costs: deliberately not optimizing, yet
On token consumption, Scanlan says the opportunity cost of optimizing is currently too high, so Intercom isn't focusing on it. They acknowledge that the spend is "an extraordinary amount of money," that they consider it appropriate for Intercom, and that the current growth rate isn't sustainable. There are known short-term wins Intercom hasn't pursued. Scanlan expects harnesses to improve, citing recent Anthropic features around model selection and anticipating better caching, but still expects spend to keep rising. The current priority is avoiding sloppiness: no over-complex orchestration that burns tokens for fun, no Opus where Sonnet would do. Playwright is one area Scanlan says burns many tokens and can be improved with minor changes. More attention to cost is expected "in a few months' time."
Robby asks how teams with skeptical finance departments should handle this. Scanlan admits Intercom is "easy mode": leadership is all in, and the finance team are heavy Claude Code users themselves, which Scanlan says reduced the questions. From years of managing Intercom's AWS account, its biggest spend, Scanlan had regular conversations with finance. About a year earlier, when finance asked when the growth would end, Scanlan sent them a Steve Yegge article, which Scanlan recalls as "Death of the Junior Developer," predicting token spend would grow 10–100x within a year or two. Scanlan's broader framing is that AI spend should come out of the headcount budget, at a similar scale, rather than being treated like tool seats such as JetBrains licenses. Scanlan argues that for most businesses, and certainly progressive tech companies, running as fast as possible is critical. If budget holders block adoption, that should be escalated to the C-level immediately, because companies that move conservatively now "are going to look pretty bad in a year or two."
Ephemeral code
Robby floats an idea: an agent might write temporary classes, overlay them on the Rails app, answer a question, and discard the code, keeping the codebase and context window smaller. Scanlan finds it interesting and points to Intercom's rake tasks, which are largely one-off code that has to live somewhere so it can be run and reviewed. Truly one-off work might end up flowing through the Rails console MCP instead of Git. Scanlan stresses that an audit trail is always needed, and says the MCP would probably need more state and helpers to serve as that kind of execution environment.
Does Rails hold up in the AI era?
Robby asks whether Ruby's smaller share of training data hurts. Scanlan says that two years ago they would have said yes. React, Node, and Python got better results than Intercom's Ember or Rails, especially the custom Rails that accumulates in a large codebase. Something changed last year, and Scanlan attributes it not only to models but to harnesses getting better at discovering and navigating large codebases. Claude Code started getting Intercom's Rails code right the first time.
Scanlan mentions a recent study making a case for Ruby being LLM-friendly, but mostly relies on a simple heuristic: what's good for humans is good for LLMs. That includes docs, tests, short code, and Ruby and Rails' expressiveness. Perhaps LLMs once favored verbose, Java-style boilerplate, but Scanlan says they no longer see Claude taking wrong approaches in Intercom's codebase. Asking it to make code more idiomatic works, and its Ruby doesn't read like a Python or Java developer's. Scanlan hedges that they haven't used many other Rails codebases, and that unusual structure or something like Sorbet might be a barrier. For Intercom, "the teething pains" of two years ago are over.
Wishlist: an agent-first Rails ecosystem
Looking 6–12 months ahead, Scanlan wants the Rails community to treat agents as the primary users of the framework: choosing it, installing it, operating it, and querying it for telemetry. Scanlan asks why LLMs don't pick Rails by default. They mention a podcast in which John Collison described finishing a side project before thinking to ask what language it was in. Scanlan says agents currently tend to default to Node or React, and suggests the community work on the zero-to-one experience of installation, setup, and documentation, a kind of "agent SEO" or agent experience. Intercom is doing similar work to make its own product appealing to agents.
Concretely, Scanlan would like not to have to build something like the console API themselves. They want native or well-supported community gems for agent access, with Console1984-style safety and good defaults, so that choosing Rails, or having your agent choose it, gives you the best agent-first environment. Scanlan also suggests revisiting some conventions on the assumption that agents, not people, are driving, possibly discarding conventions that don't serve agents. Agents take the path of least resistance, so if configuring or enabling something in Rails is easiest, they'll use it.
Asked how to get feedback from agents, Scanlan suggests writing evals to find breaking points and benchmark over time. You can ask agents or dig into session data, but Scanlan says they aren't self-aware and will produce plausible explanations that may not reflect what happened. Practical steps include serving developer docs as markdown and making search accessible, so agents reach answers in fewer steps. Scanlan compares this to latency in shopping funnels: more tokens and steps lower the "conversion rate," even if the agent eventually gets there. Scanlan believes Ruby on Rails is already very LLM- and agent-friendly, and says the community should do the rest of the work so agents adopt it.
Closing: re-reading the classics
For a book recommendation, Scanlan bends the rules and picks a technical book: Martin Kleppmann's Designing Data-Intensive Applications, recently reread, possibly prompted by a new edition. Scanlan's point matches the earlier nod to McKinley. You don't need to read every book, just know a handful of core works well and keep revisiting them. Scanlan didn't learn anything specific from this rereading, but was reminded to keep the fundamentals top of mind in day-to-day work. That echoes the pilot metaphor: even as more work moves to the autopilot, the humans running the system need to stay fresh on the craft.
Welcome to On Rails, the podcast where we dig into the technical decisions behind building and maintaining production Ruby on Rails apps. And I'm your host, Robby Russell. I run Planet Argon, and for over 20 years we've helped teams maintain and evolve long-lived Rails apps. So I tend to approach these conversations through that lens.
In this episode, I'm joined by Brian Scanlan, who's a senior principal engineer at Intercom. Brian's been working on something pretty wild lately. Turning Claude Code into a full stack engineering platform inside of Intercom. We're talking about a system with over 100 internal skills, hooks that enforce engineering workflows, and even read-only access to production Rails data. Used not just by engineers, but for product managers, support, and design teams.
In our conversation, we also dig into what happens when your CI pipeline starts to melt, and how that led to rethinking the entire development life cycle. And what it actually takes to build world-class AI-assisted workflows on top of any existing Rails codebase. Brian is from Dublin, Ireland, but happened to join us from Intercom's San Francisco offices. All right, check your belongings, all aboard.
Brian Scanlan, welcome to On Rails.
Cheers, Robby. It's great to be here.
So I have to start off the conversation like I ask all these conversations is, Brian, what keeps you on Rails?
Sure. Well, Intercom is a 15-year-old B2B SaaS with millions of lines of code. And so, you know, most of this is in the Rails app and it's a lot of effort to break it up into microservices. You know, obviously this would be something we would love to do. I joke. You know, we love the expressiveness, the ease of use, and just the normal reasons why people love Ruby on Rails in that like it's so easy to get started. It's kind of meeting you where you are and not like through a lot of boilerplate. And it allows our developers to really focus on building product and not learn gigantic frameworks, build complex microservices, and things like that.
You know, we think the Rails philosophy and like having a single application where we put most of our logic, we think it's very consistent with our approach to engineering in general. We're a principles organization and what has worked well for us over the years has been to be technically conservative. And the way that this comes out in practice is that we end up with large applications where we eject all of the smart stuff that we put in place to make sure that people are as productive as possible and they have to think as little as possible on all of the gunge and undifferentiated heavy lifting, and they can spend their time considering and shipping great product.
I like that. There's your case study for Ruby on Rails there. And I don't think I've heard that technically conservative. Yeah, I like that framing. Where do you distinguish that between like say boring?
Oh, I love boring. Yeah. Dan McKinley, he's ex-Etsy, he wrote this blog post a good while ago at this point called Choose Boring Technology. And that was a deeply influential piece of writing. He's also written a bunch of other great stuff that I refer back to constantly. It's like one of the blog posts or series of blog posts that every so often I just remind myself, I need to read that as a refresher. It's like this stuff is gold.
But yeah, Choose Boring Technology came out of Dan McKinley's experience at Etsy where he just builds things on top of a bunch of boring, well-understood components. You don't really need to think about it after that point in that somebody else will probably scale it. The work will be absorbed as part of the greater effort to say keep the MySQL layer running well or the Rails app running well. Whereas, if you're on a team, you're making a choice around what technologies to use, and you're using some pretty cool, bespoke, custom, or unique technology that's really, really well suited to the problem at hand, then, you know, it's like owning a pet at that point. You now have to maintain that thing and feed it, and like it's not the worst approach in the world as in it's better than doing nothing, but like I think that in most businesses sticking with a well-understood set of technologies that you invest in deeply and understand them well, that just gives everybody back like more throughput, more time to build product.
But you get to do things at a higher quality as well. It's hard enough to scale one database. I can't think of how hard it would be to scale three different databases if all your teams kind of picked different databases. So the way we articulate that in our environment is being technically conservative. It's like the articulation of what we have done that has served us so well in the last 15 years of Intercom.
Before we go a lot deeper into some of the topics I wanted to have you on for, for our listeners that might not know much about your role or at Intercom, what are you actually responsible for day-to-day within your role there at Intercom?
Sure. Well, these days it just feels like I'm some sort of Claude Code evangelist or whisperer. But so, I've been at Intercom for 12 years. I work on our platform group. In practice, this means that I care about Intercom's availability, performance, cost, security, and developer productivity. So everything that it takes for Intercom to be online, to be built well, to be secure, cost-efficient, all of that. And so I work across all parts of engineering and beyond to help out, you know? I'm happy to do everything from on-call to interviews to high-level architecture work or whatever.
But yeah, because developer productivity falls under my teams and we take care of our monoliths and we take care of our developer environments. And I guess my team, my group, they're really strong culture carriers inside of Intercom and we're quite influential over the last couple of years when it's come to the adoption and rollout of AI in Intercom, which we are very bullish on and making good progress. Me and my team have been really pushing on this and doing a huge amount of enablement work, research work, and getting results. It's working out well for us so far. Though we're still incredibly impatient and really, really want a lot more out of the system. But say yeah, we're having good—
Sure. I like that. That tracks. I think just for our listeners a little bit more like who might not have an opportunity to work at an organization that has a platform team or a team that's specifically focused on those challenges versus being like, we're building the tool, we're building our app, we're working on our app. Or if we want to improve it, we're the ones that have to do it cuz we're the ones that live and breathe in it every day. And like is it safe to assume that like the platform team is also responsible for figuring out like Rails upgrades or who owns these types of things there?
So we definitely own Rails upgrades and beyond. We consider like the point of handoff between ourselves and our product teams who build inside of the app to be well certainly a lot deeper than just the current version of Rails, but all the way to say the interfaces that they're working with, whether it's at the architecture and implementation level. So if it's asynchronous workers or our web servers or all of the glue, the kind of patterns that need to evolve and emerge to do that, you know, my teams own the observability, the design implementation, features, just everything.
And so, we have I guess hundreds of engineers building on top of this software, and we make sure that the Rails app and a few other apps as well, we have a bunch of front-end stuff as well, is all architected to suit and be optimized for these guys. So like we take care of everything from the AWS accounts all the way up to writing and owning a lot of Rails code in order to support the application. And then we partner and pair with them on any kind of operational issue or database upgrade, database optimization, query optimization, all of this. And we do a lot of proactive work in the area. So it's not a case of like we treat our Rails app as like a black box and people come to us with their problems. We own the problems and we will in some cases change how the application works or our features work in order to make sure that like we are cost-effective or performant or that we're unblocking upgrades, whatever is necessary.
Is there a concept of like a DBA type of role there at... What you're talking about the platform reminds me of like my early era of my career where there were a couple of people that let's say owned the database. And like if we wanted to add new features, we add new columns or whatever, we had to kind of negotiate that with them. And they would be like, all right, well here's what you actually can have. Here's how you can integrate with it. Here's the stored procedures or things like that. If someone's building out new functionality in the Rails app, how much collaboration does that team have with your team to plan that out? Or are they able to just build their typical Rails thing and you'll come back and I'm going to say retroactively optimize that, but how's that kind of work there?
Yeah, so we want to make it as seamless as possible. And so that people can just use regular migration mechanisms, regular Rails mechanisms in order to get their code working in production as easily as they can in a more simple environment. The way we go about this is that we don't staff our team with system admins or DBAs or anything. And this is a little ironic cuz I probably identify as like a system admin, SRE in like in my own kind of background. But we are typically staffed with product engineers. The usual way that people join our teams is that they would be a regular product engineer building customer-facing features inside of Intercom, and they start to get into like some niches or they start to like do a bit of work. They might end up doing some database performance work or whatever. And they'll also see how great my teams are and go like, I want to work with those people. As in I want to change my career or change the direction I'm working in to go like say deeper in these areas, whether it's deployments or databases or Rails in particular.
And so that's how we recruit and that like don't hire in specialist knowledge. And this is true across all of our product engineering roles. We don't tend to hire Rails people for example or look just for product engineers with 10 years Ruby on Rails background and experience. Now for sure we definitely need a decent bench of people who have got this depth and knowledge and experience, but we find that hiring generalists who can go deep in particular areas, willing to learn, willing to throw themselves into those areas works better both from like a generalized hiring into the product engineering role and then also into the platform groups. I think getting people who have wants to build for themselves as such and want to go deeper into those areas, but are, you know, legit software developers who understand Intercom and get stuff done. That's the ideal person to join our groups.
Reflecting on your role there, what types of things keep you up in your role up at night?
Well, today it's we're not moving fast enough on AI adoption, which
Of all companies you're not moving fast enough relative...
Like yeah, relatively speaking we're not bad and I talk to a lot of other companies and yeah, I'd say we're probably up in the very high percentiles, but it's not good enough and it's like we have a lot of anxiety that say companies we're competing against, they won't have the same headwinds. For all that I love about the Ruby on Rails monolith with all its code and tests and stuff like that, if you're working on a greenfield company or greenfield environment, you haven't got all the baggage that we have. All of the kind of complex features that we built over 15 years, all of the different styles of code in our code base. Well, they might have revenue, but like revenue is good, but the production of code and production of features and our ability to do that incredibly fast and accurately, that's my job or like part of my job and we cannot afford to be held back by, oh, we've got a million lines of code, therefore it takes Claude Code five times as long to kind of get at something that's good or whatever. Like it's kind of unacceptable.
We're paranoid about the headwinds. We also know that, you know, if the tools don't improve at all, like we're on a path to basically assembling kind of Voltron-style a lot of Claude Code skills and guidance and stuff like that that can do the job of like the vast majority of things that our product engineer does in Intercom. And so we're impatient to get like we need to finish this. There's a lot of work to do. Turns out that humans do a lot of work. It's like trying to move as fast as possible to get as much coverage, but maintain a super high quality bar with the kind of the transformation to AI because I'm not that worried that we're going to go down too many paths. I think you can notice when things are going bad, but I want to get this stuff right first time and so we're on a big transformation. It's a lot of work changing out the corners of people's work. It's a huge amount of effort.
Yep. So for some context for our listeners, you know, prior to this conversation Brian recently wrote an article. I forget like... One of the things I didn't write down was the actual name of the article, but what was it off the top of my head?
It was just a tweet thread. Yeah, it went down to a few different places. I think it got called how we use Claude Code today at Intercom.
Anyways, I read that and you know, and like I was like I need to talk to Brian again. It's been a while since we chatted. I had him on my other podcast a couple years ago, but the thing that stood out to me was how much pressure your CI system was under.
I know.
To the point where it was like maybe starting to break down. So maybe you can walk us through what was going on there.
Yeah, so this is one of those success stories. Very often like a lot of my job is unblocking people or giving them permission to do things and I'll say something like, oh, if that problem happens, you know, some problem of success or saturation or scale, well clearly we'll have been wildly successful and the first thing we should do is throw a big party. And then we figure out how to fix the problem.
And yeah, in this case the wild success was that we were on a rapidly increasing rate of throughput in terms of pull requests going through our system. We have a large test suite for our Rails monolith in the order of hundreds of thousands of tests and it was taking about two and a half days of wall clock time to run all of the tests. Now we highly parallelize these things and we make hundreds of changes a day. So we're pushing large numbers, running millions and millions of tests in order to ship up to 100 times a day to production and this is the entire way we work. Like if we can't ship, we're down to a certain extent. Like we treat it as a P1 event. If it's a flaky test, if it's something stuck in the pipelines or whatever, it's just we are so addicted to this stuff that once it's removed, it's very noticeable if we can't ship at any moment in time.
And two things started happening. One was we started getting a lot of flaky tests in the environment. So we had rolled out Rotoscope in kind of the previous year or so and so Shopify's Rotoscope, what we use it for is to identify changes where we can run a subset of tests and not the entire test suite. So when you have a test suite the size of ours, this buys you back a huge amount of compute and potential cost optimization for running all these tests. It doesn't buy you too much in terms of the time it takes because it's always going to be your slowest test that will kind of dictate
that. There's a lot of setup and teardown. But so this increased velocity and this use of Rotoscope just started causing more and more retries, more and more flaky tests. So if any test fails in our pipeline, we just automatically retry it. The way things will go is like if something is legitimately flaky is it will fail with one seed or one order, in one order, but it will be fine next time and we're shipping all day. This is kind of fine. You just ship it again like five minutes later, it'll be fine because that test will have run again and the seed will be different and then the order will be different and your code gets saved. You get to some sort of critical mass of where you get this small number of flaky tests and then suddenly they're tripping over each other multiple times and this is bad at this point. Additionally, we saw like a non-linear or maybe like linear but growing faster than the rate of throughput of the cost of this system.
Yeah. And when I say like CI blew up, it's mostly like the cost was growing like faster than our business. And our business is growing fast, so this is bad. And so obviously we didn't turn it off, you know, we were still shipping hundreds of times a day, but we were uncomfortable with the rate of growth and the amount of failed tests through it.
How do we fix this? We just got down to a classic kind of tuning fix-up kind of cycle. We resized the hosts that we were running on it. Like basically everything was in scope. We looked at everything from, okay, where are we running these things? Should we remove spot hosts to buy stability, because we use aggressive use of spot hosts. A spot host is a host that you can get from Amazon. It is like preemptible, or rather it can be removed at any moment in time, but for the privilege of this you pay a lot less money. And so Amazon get to benefit from being able to do dynamic capacity management and so they're kind of selling off like spare hardware. It's not quite spare, but because they can take it back, they were able to make money out of it in different ways. So it helps their margins and it helps us because you can just get compute, and for things like tests, you know, it's very appropriate to run these on hardware because hardware can go away and you just rerun the test, but you still have to tune things to work it well. If a host is signaling that it's about to shut down, you don't want to pick up new tests and stuff like that. So there's just a matter of tuning and stuff like that that you need to do in these cases, and so like we stopped using spots for a bit just to kind of stabilize things, keep things working well.
And then it was a case of like looking at every single bit of performance and squeeze out what we could. There was no like one big change. It's like, oh, you just need to do this one magic thing and then your tests will be in great shape. We applied the classic boring scientific method. Like use observability to understand where the bottlenecks are. Turns out it's mostly for us use in loads of factories. And we have all these big kind of semi-god objects in Intercom like conversations and workspaces and these very common things that had a lot of heavyweight setup that was being invoked hundreds of thousands of times across our test suite. And it just turned out that you could do cheaper ways of doing a bunch of these things, or we had, say, a very deep conversation object that was able to be used on virtually any kind of test. And it's like, well, we could actually just make a lightweight one. [laughter] And only do all of the complex stuff when it's kind of needed.
And so again, not kind of rocket science, but you just need to be able to get that kind of feedback loop of like, okay, let's identify the places where the time is actually being spent the most. We're able to use Honeycomb to have like deep tracing information that shows us this, and then pretty much go after the sources of like the largest amount of time and flakiness, which tends to be in the kind of factory type areas. And once you know where it is, you know, you've got options and you can start to attack the problem in different ways. But it wasn't one of these. It was like dozens and dozens of these fixes that then contributed to the CI system being pretty much in a better state right now.
You know, kind of circling back a little bit to the high increase in PR throughput from your team through Claude Code. Can you tell us a little bit about what changes your team started to make on the engineering side? I know that we can talk, I'm looking forward to talking about how people outside of your engineering team are using it, but for working on the development of the app rather than just focusing on the, you know, using some LLM tooling within your product, but like actually on the development life cycle, what sorts of changes did you start to make? And then for someone else listening, if they're going to start doing this and they haven't already and they're kind of curious about it, do you feel like they should make sure that they feel pretty confident about their CI workflow process and optimize that a little bit first? Because like almost everybody that I've talked to has talked about like, "Oh, wow, we didn't really expect how much our GitHub Actions were going to start blowing up in cost." And like, "Oh, we're running out of budget here." It's great, higher throughput, but it's either humans are the blockage because there's not enough time to review everything, or they haven't figured that part of the process out. But I'm assuming you just didn't start doing this in like January over at Intercom, but tell us [snorts] a little bit about that progression there.
So, we started seeing the increase in throughput around last summer. That's when things started to kind of increase for us. It also coincides with when our CTO, Dara, published a goal of doubling the throughput of our engineering team. And really, the measurement that we're using is pull requests. And so, pull requests per head, per number of people in our R&D org. We've got the expectation that, you know, when you adopt AI tooling you should have more time to work on code. You should be able to produce more code than ever before and get more stuff done. So, there should be a very natural rise in this crude measurement of pull request throughput. Of course, look, every measurement isn't a measure once it's measured and all that. But I think it's still like a reasonable thing to expect, that like life should get so much easier when it comes to the production of code that there'll be a natural increase through this.
We're working on our CI system. We knew that there was low-hanging fruit there to kind of go after anyway. So it wasn't surprising that it started getting into trouble. Other areas that we started working on like, you know, we already had a pretty good linting system, a decent amount of RuboCop and some other checks in GitHub Actions. But we started doing more and more. Just like getting more and more bullish, or like again kind of maybe looking a little down the line, but like, you know, as you're seeing more code being produced, you want it to be done just better. And Claude Code, God bless it, like it still can get stuff wrong. And without a decent set of linting guardrails, you know, it can just start to produce code that's confusing, or code that's inspired by code we wrote 15 years ago, which is even worse. And so yeah, if we zoom out and just look at the amount of cops that we have in place, and even the sophistication of the cops that we have in place, like they're just way better now. Claude Code is really good at writing decent cops. And you know, the syntax is a bit weird. It can be a little bit intimidating. And it's kind of like the counterfactual is that we never wrote these and then we end up with a massive code or something, but I'm pretty happy that we sort of started to see that this was coming down the line.
Other things we did, but with like mixed success, was at one stage I counted and we had like 10 HTTP client wrappers in our Rails monolith. You know, for whatever reason, people would like wrap individual service calls, or we just have like HTTP clients and Typhoeus and, you know, all of these kind of peppered around the place. And so, we started doing a bit of like consolidation of those things, kind of with the idea that like we don't want to be confusing the poor LLMs. I guess I framed this as like, "Look, this is good for humans." And so, if it's good for humans, it's good for LLMs. Like the agents just need to know the one way to do things. This work kind of petered out. It was hard to kind of prove that it was making a big impact. It's real soft kind of stuff. And I think it's easier just to configure like guidance to, say, Claude Code or whatever and just say, "Look, if you're going to be writing an HTTP client, use this HTTP client," rather than do a lot of surgery across the app. So we've kind of ended up leaving in place a bunch of patterns, but then just telling Claude Code like, "Hey, do not use metaprogramming," and stuff like that. [laughter] And you know, it's not going to accidentally pick it up from the code base. So, being direct there was just more feasible than boiling the ocean on fixing the entire code base.
I just want to kind of drill down a little bit deeper here. Like, let's say, approximately, we're recording this in the middle of April 2026, for listeners. And I'm not sure exactly when this will get published, but right now let's say you have a, I'm not sure where your team manages, like, what comes next in the backlog of product. There's a feature or change to be had. What's that workflow look like right now? Is it like, do you have a Claude skill? Like, an engineer gets a ticket assigned to them, or they select a ticket off the backlog, and then like, "All right, work on this, Claude." And then just assign the number to it and then it does the bulk of the work. Are they sitting in planning mode in Claude Code? You've mentioned Claude, so I'm assuming it's kind of standardized at this point. Maybe we can dig into the pros and cons there, but like what is it at that very, very basic level? Like, I have something to do today. Where does Claude come into the workflow right now?
Sure. There's a good few points here. One is that we've mandated that Claude Code is our tool. And the important point here is that you pick one, because without one kind of platform, it's very hard to kind of move up the level of complexity or sophistication to build repeatable skills to do things like all of the work that we do, whether it's Rails upgrades or fixing a flaky test or responding to an outage or writing brand new features. Like, all of this must be agent first. Greg Brockman had a blog post or Twitter post, whatever, earlier on in the year, and he stated that by March 31st, all technical work in OpenAI was going to be agent first. And I think he's a little bit ahead of us, but we have the exact same internal principle. It's like all technical work is becoming agent first. So, if you're not opening an agent first to do that piece of work, whether it's responding to an outage or doing, say, some planning or research or starting to write code, then you're kind of doing it wrong.
So, now there's kind of different levels of maturity in the way that people will work with agents to get something done. And you know, just the capability of an agent, like what can you actually do at the moment. Where we're at right now is like well over 90%, 95. It can be even higher some days. That's the percentage of code in our Rails code base that is changed by code generated by Claude Code every day.
We want people to not be reading the code either. Like, I don't really open up my editor these days. We want people to like shut down their IDEs. Just tell Claude your problems and let it figure out which skills to invoke and produce the right code first time. That is the vision, and we do see a lot of that. Now, a lot of people are still with editors. They might even be invoking Claude Code from inside VS Code, or they might look at the code after it gets invoked or whatever. But I think that like those levels of maturity, where at the start you're kind of using tab complete, and then you're getting it to write most code, then you're just not opening up your IDE, you're not really looking at code, I think most people at Intercom are starting to get pretty far down that journey.
We're abstracting and guiding people away from worrying about the code, and their work should be plan mode style stuff or whatever. But, you know, it doesn't have to be a big spec. I think there's a lot of shadow of it, like spec-driven development and people really tuning their kind of initial stuff. I think like chats, conversations, interviews, they're all good ways that you can invoke the agents to work with you. And so, I tend to not produce these giant specs and get it to do the right thing. It's more like interactive and kind of figuring it out as it goes along.
And where we're not universal on this is like there's still lots of work that is artisan. The way we're thinking of, say, the software factory is like if you want to produce IKEA furniture, like this is the software factory. And this is what all the kind of skills in Claude Code are, what we're trying to build and assemble. But if you're doing like pretty ornate, unique kind of work, maybe the factory's never seen this stuff before, or it's something new, like it's of a higher quality or it's got a particular quality where it's not appropriate to build a factory around it, then sure, we're going to have skilled artisans kind of producing this stuff. And I think like UI as well, like front-end type work, Claude Code still isn't good enough at like detecting, "Ah, that's off by two pixels," or whatever. I'm sure it'll get there, but there's still a lot of the kind of more visual work which can be aided quite well with a human in the loop and an editor in the loop.
So, yeah, there's a maturity level that we want everyone to get to. The vast majority of times people shouldn't be looking at code. We want that code produced to be pretty perfect first time round. And we're starting to see a lot of that. Our job, like my job at the moment, is just fill in the gaps. Whenever that doesn't happen, figure it out, find out why, fix it.
So, to make sure I understand, with the agent-first approach, is there still like a human assigning a thing to an agent? Like in their Claude Code CLI, is that the kind of UI most of your team is working on then? The CLI? Or do you have some other machines running these independent agent Claude bots, or they're fake people that, you know, you've given them a persona and they're like, "All right, you're this experienced developer. Just pull tickets off the backlog and work on them." Where are you kind of falling into that camp right now?
Yeah, today mostly it's driven from local Claude Code.
So, right now, if I'm paged into an incident, I'll just habitually open Claude, tell it, "Hey, I'm in this incident." Maybe I should automate that. That's something I might do. And then it'll join and figure things out or whatever. We are working on remote agents. There's some really good stuff published by the likes of Stripe and Ramp and others. And there's a few startups kind of building these things. Anthropic are doing their own thing and all that. There's plenty of options, but we're building our own internal agent system. We know what we want to run. It's effectively mimicking the Claude Code setup we have locally. But right now, where we're at, we're getting so much value from just people locally on local agents and kind of driving them manually. And like, there's been some good features recently around being able to run scheduled tasks and things like that.
where the competitive advantage of being remote is certainly like reducing it. But for, I guess, multiplayer or long-running or large tasks or scheduled tasks or tasks that need to be brought in through automation, yeah, you got to have remote agents that you can automatically get to join incidents or join Slack channels or whatever. So, that's one of the things that we're working on.
Interesting. But you can just get a lot done just like without any of that fancy stuff. Just local Claude Code is a perfectly good way to get through a lot of work.
We've encountered this ourselves as well with some of our clients is that individual developers are using the tools, they're submitting PRs, and they're still working through a lot of the same steps they've been working out, but they're finding like, "Oh, now the amount of time that I'm spending having to review PRs is maybe higher because there's more PRs coming through."
And then they start to ask questions like, "Should humans be reviewing the PRs?" And then, "What's the purpose of a PR if the bot, you know, or the agent created a thing and it supposedly followed my linting rules or my RuboCop rules?" How is your team thinking about those types of things right now? Are you removing any steps that used to require humans? Or do you still require, like, "Hey, PRs need to be reviewed by someone else, but it could just be their instance of Claude Code." And then, that way you had two people doing it, but if you're not looking at the code, what are the humans actually doing in those steps outside of assigning the work to be done?
Yeah, so, code review is the current bottleneck. And so, this is kind of like the purpose of increasing throughput through the system. You increase throughput, you find the next bottleneck, you fix the next bottleneck. And today we are aggressively going after — our goal is for the majority of pull requests in Intercom to be approved entirely automatically and not have a human approving the pull request.
Now, we're not reckless. And we've been doing similar work for quite some time. So, this is a journey. And we've been doing some kind of novel stuff or some interesting things to get there. But today, something like 20% of pull requests for our Rails app are completely approved automatically by our system that uses Claude Code to automatically approve what we call safe PRs.
Now, let's wind back the clock a little. Many years ago, you know, GitHub produced Dependabot, and I love it. It's great. And we got it to automatically create pull requests for, you know, security updates to different packages and gems, and great. And we would automate this, and we would open up an issue for a team, and the team would get the issue, and then they'd ignore us. And so, we built up like hundreds or thousands of these automatic pull requests for — yeah, I mean, I don't think I'm unique here describing this. And you know, we've got a long tail of applications as well, and all of this stuff just builds up.
But the other thing is, what are we expecting people to do? Like, if I'm upgrading the, I don't know, the JSON gem or something. Am I really going through every single — like, what other thing is running the test suite and pushing merge? And like, what value am I actually producing? Especially with like a super large code base. Like, we just have to rely that all of our unit tests and deeper integration tests and smoke tests that we run and post-deploy exception monitoring and everything — we just got to assume that all of that is perfect. And that is working to catch any issues that an upgrade of these libraries would do.
And so, yeah, a few years ago we stopped having humans in the loop in these and just automatically merge GitHub Dependabot-created PRs. Now, we do it at a certain time of the day. We do it like Tuesday, Wednesday, Thursdays. We do them in batches. You can see them in a Slack channel. So, it's not like they're happening at 2:00 in the morning. And we opt out a few different gems, you know, the AWS gem, and in our Rails apps. I just don't do that. Don't do Rails, you know. There's a handful of things that are like LangChain in Python as well. We've just found through experience that you want a human in the loop for like a small number of load-bearing dependencies.
When I describe this to people though, like outside of Intercom, they're kind of like aghast that we would trust the CI system so much and like not have a human in the loop and stuff. But I think the main thing is to look at like, why are we having a human in the loop in these things? The value is just incredibly low. We are way better off encoding the important stuff. There's times when you do want a human in the loop with these things. And like, this is a thing that's respectful for our time, you know? It's like, it's one thing just to do work because it's a process and because it's the way things are done or whatever. But we just don't accept that people should be doing this kind of low-impact work.
And we have a bunch of other deterministic rules-based automatic approvals. If you're updating a spec in our Rails monolith, who cares? Ship it to production. Like, you don't need approval for that. If you're changing some documentation, some other kind of bits, you know, different things where we've kind of used path-based approvals, we're happy to give automatic approvals. And this stuff has been encoded in our like SOC 2 and ISO 27001, all of our compliance stuff. Like, these compliance regimes, despite rumors to the contrary, you know, do not involve humans having to approve each other's work. They require you to write down what you do, have a bunch of controls and mitigations and risk reduction mechanisms and stuff, audit trails. All of these kind of good things. These are completely achievable with automatic review processes. In fact, there's no reason to think that we can't build better review processes with a mix of agents, rules, and like remove humans from it, and you get like a lot faster and more reliable or repeatable kind of processes.
So, for the last while we've been tweaking and tuning and trying to get good feedback into pull requests from, say, our Rails monolith. Initially, you know, we turn on Claude Code and say, "Hey, review this code." And it would produce a bunch of reviews. The reviews were terrible. And they looked like good reviews. And this is the kind of hard part here: agents, LLMs, they're really good at producing stuff that looks like real good work. But they're very eager to say anything. And they will also be a bit psychopathic at different times, and they'll make a big fuss out of things.
And so, the first thing we kind of ended up doing was like, "We got to put manners on these bots. We got to tell them, do not get involved unless something is very impactful. Keep it short. Also, if you've got nothing to say, say nothing." And then also we gave it different examples of like, "These are good examples. Like, here are the kind of things that we want you to go after." And just started running that. And like, the great thing is that we've got so many examples of changes gone through the system. It's just not hard for us to be able to produce a large number of like potential reviews. And then we compared them to the reviews done by like our best Rails engineers. And then we got our Rails engineers to grade the output from the LLMs.
And so, we used the data and the expertise that we have and the LLMs to like build a bit of a feedback loop to make sure that like, okay, the output of this stuff, we're not just one-shotting a bunch of reviews. We're being really picky and careful about trying to produce like a super high-quality reviewer out of this. Something that like we'd be proud of actually producing. We tuned this reasonably well, and then people will start actually taking action as a result of reviews. If you've just got noisy stuff, which is reviewing, showing up on everything, pointing out all sorts of stuff that's not actually that important, people just ignore the code review feedback. But by putting manners on them, by paying attention to what's important to us, and by having us in the loop, we've got a feedback loop, and we started getting pretty good code review feedback.
And then the next step is approvals. And so, it's pretty much the same process again. It's like, get your best people, backtest all of your old reviews and approvals, and figure out what kind of criteria you're happy with in your environment. And for us, it's a mix of historical data, information that we got out of our experts, but also just a bunch of things that we think are sensible. So, I think like the perfect change isn't hundreds of lines long, and it's behind a feature flag if possible, and you know, it doesn't touch, say, certain code paths, and there's all these kind of qualities that when you look at a piece of code, it's like your question is like, is this safe? Is it safe to approve? And so, we've encoded those so that an agent can figure out, is this safe? And if it's safe, we're pretty good with getting it into production as quickly as possible and unblocking.
Of course, there's other things that code reviews are useful for. They're a teaching aid. They can prevent all sorts of other problems that come down the line. But ultimately, we rely, or we're betting on, us being able to put pressure on the system, get more throughput, learn from anything that does go wrong, and improve our like checks and balances and code base or whatever, or maybe even come up with new ways of deploying. Like our deployment system is very, very simple. And maybe we need to add more features to make it easier here, but like it's going well for us, but we want even more. We want to get like over 50% approved.
Yeah, yeah. It's hard for me not to just like — if you're already going through the process of like, individual engineer on your product team collaborates with Claude, ships a PR, then someone else has one of their agents review the PR. And it's like at what point is the multiple people in that process really useful if they're not really looking at needing — if their desired outcome is like we don't even need to look at the code. So, why not just have the engineer run a series of different skills. It's like, all right, submit the PR, new fresh context window, review the PR, fix PR stuff, merge. And then there's no other person involved in that workflow. Like why do you think it's still valuable right now to still have those kind of checks with other people? Cuz I think we talk about accountability, but the engineers are accountable for the stuff they get shipped, but they're also like, but don't look at the code. We'll just rely on Claude to do it all. So, it feels a little bit like mixed messaging in a weird way. So, I'm not accusing you of that, but I feel like we're weighing up these two kind of like different spectrums at the same time and I'm trying to make sense of it currently right now in my own professional career.
Yeah, it's tough. Like there's definitely stuff that goes on in the code base or, say, our infrastructure that I want to know about. And I think code reviews were kind of the default way of like having this kind of gatekeeping or blocking mechanism. And it's a good way of like getting visibility into changes that are going on and stuff. I just think we can do a lot more of that reactively and also with agents. So it's just not that hard these days, say on a weekly basis, to say, let us know of any changes. Let's say you're working on the Intercom Messenger, you're on the Messenger team, and like just give me a bunch of changes, or you know, you can get Claude or whatever to say, give me a summary of anything that's gone on in the last week. And that will kind of keep you up to date, and that can solve the same problem of where a team are working in an area. They want to know what's happening.
But I think the act of blocking people from shipping is a pretty extreme solution to the kind of problems of like not having, say, specific concerns that can be resolved through actually having the agents at the feedback level or at the code review level say like, no, no, we do not touch that piece of code, or rather you have to get a human involved because that is the busiest code path in Intercom or hottest code path. It's like that's worth having the ejector seats go out and just like, yeah, just get a human involved. So I'm okay with like specific areas like that. And then there's areas like education and, yeah, understanding what's kind of going on in the environment. They're not blocking. They're things you can do reactively, like after the fact that the code has changed. Yeah, I think there'll always be code and things that, again, like artisans are kind of writing in your environment and that you'll still want oversight in those places. So I'm not saying that like code reviews serve no purpose or that like a lot of the functions that were going on were invaluable, but the main thing is like most code is just not that interesting.
And especially if it's been through like the heightened amount of linting, tests, everything, all the stuff that we're doing. And we're kind of forcing it into this auto-approval mechanism, which forces the code to be like a small change with tests and blah blah blah. It's like you can really take a little bit more risk, or what would have been riskier in the past is like net far less risky now. Like for example, we're working on a deployment agent. Like something that has an agent monitor 100% of deploys and do like pretty deep, like looking at the code, looking at exactly what metrics should be exposed. The kind of stuff that like maybe if you're pushing out like a big change like a Rails upgrade, you'll have your best engineer monitor it like a hawk. But now it's like we can do this for every single change. We can have like the equivalent of our best engineer, exactly what they would do, and maybe even more, just have them automatically on every single code change. And so this is like this extra additional capability we never had in the past. And now we can do this. And again, it's like it removes some of the risk of like moving faster, or moving like twice or 10 times the amount of changes through the system and having fewer humans involved in the code review process. So, but mostly I think it just comes down to most changes are pretty boring. And just do not need that kind of blocking action.
It started like it always does. The developer said, let's just add this one gem. No ticket. No discussion. Just a quiet bundle add in the middle of the night. Years later your Gemfile is a cold case. Dozens of dependencies, no clear motive, and everyone insists it's probably still needed.
Introducing bundle detox, the Rails extension that investigates your bundle and tells you which gems are actually used and which ones are just hanging around. It tracks real runtime usage. It follows the requires. It asks the hard questions. And when a gem can't prove where it was on the night of the deploy, bundle detox removes it. Bundle detox, because in every Rails app, the call is coming from inside the lock file. May reveal abandoned rake tasks, hidden monkey patches, and one dependency nobody will ever admit to adding.
Quite a big part of your post was about talking about building out internal skills in particular. And so presumably a lot of those are for your software engineers. So, I know that this isn't just experimenting. I'm assuming, yeah, like is this one really large repository or do you have a lot of repositories spinning up at this point? So, when you're thinking about building skills for your engineers, but you've also been building a bunch of skills for people that are not — I'm going to air quote — software engineers or product engineers that are able to leverage those. Let's talk a little bit about maybe first some of the
engineering skills that you might have and then, but then we can kind of pivot over to talk more about like how you're enabling other parts of the organization. My role is definitely on the product engineering side of things. And that might actually be the least interesting part.
We have over a thousand weekly users of Claude Code in Intercom. So, there's still a few people who aren't using us. Amazingly we saw this completely viral moment of where when we got Claude Code set up and a bunch of people, like non-technical people, started trying things out and they were hooked up with Snowflake and they were hooked up with a bit of help and support. They started doing the kind of work that they had already dreamt of before. Where they can just answer their own questions. They can encode into like a skill how to like get their work done.
And so we've had this just viral takeoff of Claude Code across sales, marketing, customer support, just all of these other areas. The kind of thing to learn there, the lesson there is like, yeah, picking one tool, giving access. Like we've got good controls over Snowflake data, good access to tooling. And we had a core team of people just setting them up with great MCP configs and technical support. You just need a lot of technical support for this work. And if that's not there, then you're going to have a bad time. That's the stuff that's been happening outside of R&D and it's extremely powerful. We're seeing like people just having so much fun. Like they're doing the best work of their career. They're coming out constantly in internal channels saying like this is transformative. I no longer have to wait for months for an analyst to do some query or something like that. I can just turn around stuff and iterate. Now I can understand my customer better. I can understand this analysis better or whatever. That's so so amazing to see.
And we're just using regular Claude Code for this as well. The core of all of this and like what we're doing in R&D is kind of very similar. So, we have a base plugin. So we use the plugin system in Claude Code. That base plugin has basic telemetry and safety built in. So, we hook up Claude Code to send hooks to Honeycomb. That will include the basic metadata out of every single skill call, every single session. And the reason why we put this into the Honeycomb is so that anyone can kind of explore it. If you're developing a skill or whatever, you can go in and like dig through this data, see how it's evolved and stuff like that.
We also collect every bit of session data. So, they get pushed to... So, the full Claude Code session, which is usually like a giant JSON file or something, we copy it off to S3. We like slightly anonymize them on the way. There's only so much anonymization you can do, but we just make it slightly less easy to figure out who it is. And there's a bunch of people who we don't do this, you know, if legal or finance or whatever, we probably don't want to look at their sessions. But we copy this data off to S3 for reasons of like quality control, feedback loop, figuring out what's going inside and really looking for like, you know, there's gold in there a bit, like what people are doing and how they're going about it and is this stuff working well.
And then we've got a core plugins for like general purpose developers, developer work. And so this would be everything that would govern or dictate like how we do, say, fixing a flaky spec or, you know, doing like opening a pull request. Just all of these kind of core functions, like stuff that everyone does across all of Intercom. And we have a very high bar for those skills. Those skills would all have to have evals. So, actually tests. They would have to adhere to not just the Anthropic skill kind of guide, but we have internal kind of skill guides as well as skill improvement skills. And yeah, and so like this is like the top tier. Like skills would graduate into this kind of environment. And once there, it's like we're committed to maintaining and making sure that these kind of core skills work extremely well.
Then we've had an explosion of team-specific plugins. And this makes like it easier for teams to build out skills that might be unusable for anyone else, you know? It's like really fitting their workflow or tooling or area. And I guess the quality bar can be lower. Like you kind of want to encourage these workflows. You can afford to have this more opinionated about the setup. I think in the core skills, we want it to be very unopinionated. Should be adoptable by anyone and stuff like that. But for team-specific stuff, at this point you're getting a bit more closer to like personal workflows where, you know, overfitting to your preferences is exactly what you want. But there's still like plenty of useful stuff in there. Or you might have like a load of skills that are extremely high quality and don't to the same standard, but you just don't want everyone using them.
So for example, we have a skill that does Rails upgrades. And we've been developing this over the last few months. We've been running in production the bleeding edge of Rails during a weekly upgrade. And the work there, it's kind of repetitive. It's not that fun. Like it's very important. I think it's great and awesome stuff. But like encoding as many of these individual steps into a high-quality skill that's repeatable and where we practice real continuous improvement and it's not just like a one-shot, hope for the best. So that skill would be an extremely high-quality skill, but there's only one team that we want actually upgrading Rails apps. And so we just hide that stuff away, you know? We don't want people accidentally upgrading Rails app.
And then yeah, you get into a whole super long tail of people just writing their own personal functionality or personal workflows and stuff like that. And our thinking on this is to be very liberal on like letting people build skills, letting people build marketplaces. Like get stuff built and distributed. Partially because we want people to learn about this and get started. The other is like we might have, say, three or four skills that actually overlap. Now, we're big believers in small skills. They need to be discrete and testable. There needs to be a unit to test, you know? And so monolith kind of skills that try and do many many different things, I think they're bad. You can assemble smaller skills into like monolith skill type skills, whatever. Maybe like they discover them along the way. And I think that's fine, but we want like the core things to be super small and not very opinionated.
Maybe you might have like an investigate a bug kind of skill. And maybe we're okay with, say, 10 of these existing. And then we'll try and pull out the good stuff, like look at the session data and see what bits are usable, see which bits are getting good results. So we're happy to let like that many seeds grow or many flowers grow and then kind of hopefully pick out the best stuff. So we're starting to do that. But the main thing is to get like momentum and get these things out there and then try and pick out the good stuff rather than going slow and like hoping that like people will figure out what the good stuff is.
Sure. Sure. You know, one thing that stands out, you know, as you're talking through this and what you're building over there in Intercom, is like you've got these parts of the system that kind of need to behave consistently, at least fairly, you know, every time. And then there's probably parts where you're relying on the model to interpret things a bit more flexibly. So for folks listening who might not think about it in those terms, like how would you explain that difference? Is, you know, is that like, what do they say? Deterministic and indeterminate?
Yeah. Yeah. Like I think even today in our environment, which is well set up with guidance and all this kind of stuff, Claude Code will really struggle with very open-ended kind of problems.
We had one of our most important alarms fire about three or four weeks ago. And it was at the weekend. This is our alarm that we use that tracks actually Intercom inbox usage. And it's kind of one of these classic SLOs. Like it records the user behavior. And if our customers can't reply to their customers, they're not sending replies. Like, and that goes out, so deviates outside statistical kind of prediction, then something's gone wrong. Like so the database has slowed down or maybe the internet has slowed down. In some cases even like Carnival in Brazil has caused this to kind of like go off course.
And so the engineer who was on call this weekend did what I would do, which is like, "Oh, I got paged. I'm going to tell Claude about this and Claude can kind of help work on this in parallel." And Claude Code came up with lots of very interesting ideas as to what had happened. So it was the weekend and kind of quiet, but it started like blaming changes that were made like the day before. And you know, maybe it might, but like this was a pretty short drop and a recovery. This was like, it wasn't some sort of latent bug that was only going to kick in hours later. And then it started just looking at other things and coming up with like plausible ideas. And honestly like they fooled the engineer who was on call. Or rather they sent them down like multiple different ways. And they kind of ended up settling on one of these. Like, "Yeah, maybe it was that."
And I took a look a couple of hours later and did a bit of digging. It's like, "Ah, this looks weird." It ends up being an eventual consistency problem on Datadog's metrics, which isn't the most obvious thing you would come to. Like not an obvious kind of conclusion. And so for a few minutes I was pretty happy. I was like, "I've still got this. I can outsmart the other LLMs."
And you know, these kind of novel cases, you kind of come across them once and you encode them and then maybe Claude will just check every single time and stuff. And so, you know, you just do this a number of times and you've got like 99% of your kind of edge cases kind of covered. But that whole like confidence thing, or like going down kind of wrong paths, or like when it doesn't have a particular goal in mind but it's more open-ended, I find that it can give plausible answers, but like you really have to be pretty skilled at times to go like, "Nah, I don't think that's quite right. It doesn't kind of feel right. You know, the vibes are wrong."
And this is like one example. And it's not the most important case in the world, but like this was a very experienced engineer who was working with Claude trying to solve an issue. But in isolation, even though he's both very Intercom experienced and one of our best AI engineers, to be honest, he couldn't convince it, or he was kind of happy for the LLM to kind of come to a conclusion that in the end wasn't accurate or satisfying. And you know, enough of these happen, you start to build up maybe some wrong ideas about the production environment, or did you actually solve the root cause of the problem? And for sure, humans do this all the time. But like doing this at scale or more, but also knowing if you got something more deterministic that you can really get something like incredible out of Claude Code. But then there's certain types of work which looks kind of similar, but it just really flails and can lead you down, like it is worse to be brought down a load of red herrings than it is to like simply have Claude admit, "No, I have no idea what happened."
Like how do you think about that from like a pattern recognition thing? Like if you weren't involved in that process, the... I'm not trying to pick on that one particular person, but it was like in that scenario, like if they're like, "Well, they accept like this must... this seems reasonable. I'll take the confidence of this tool doing the thing." But then if you hadn't participated in that... Is there like a... is that part of your workflow there, that like there's someone else will come in and re-review a situation? Or is it just because this particular thing smelled a little off that it felt like it was worth having someone like you spend a little bit of extra time digging in, like a second or third set of eyes into the situation?
Yeah, I'm kind of cheating here. Like I pretty much join every single incident channel. It's like, you know, it's probably not scalable in every single company, but yeah, I definitely apply like the eye of Sauron to all of our incidents. And if I don't like what I see, then I'll kind of take a look at something.
A metaphor I've been using recently is, you know, commercial airline pilots. But they mostly land their planes on autopilot. But they also know that they're going to be tested regularly in the simulator, or maybe in person as well, where they actually have to fly a plane in different scenarios and not with autopilot on. And so they know their job is like to inspect the output of the autopilot, to monitor and manage the autopilot. But they're also eager at times. Like they'll just turn it off. They'll say like, "Hey, yeah." They want to make sure that they're fresh. They're not just going to sit back and trust the autopilot. They have pride in their work. They respect the craft and skill. And they will, for whatever percentage of landings or whatever, they will take control and disengage the autopilot. But the rest of the time they're paying a lot of attention to the autopilot.
I think we'll end up in a similar world. I'm extremely bullish that Claude Code, we can get us to resolve most incidents faster than I can. And that's kind of the bar. And we're still a good bit away from it. But even when we have it doing that, like I'm still going to race it. Or, you know, disable it from taking actions and compare and contrast. Partly because it's my job to run the system. If you're a factory owner, you need to really deeply understand the output. And you know, take the IKEA furniture home and use it and assemble it and all that. I don't think it's okay just to go like, "Yeah, it looks okay. It looks like it's doing its work." You got to be pretty active in that.
And so, that means that, yeah, even if we've got something nailed, like I would call maybe our Rails upgrades or fixing flaky tests. There's all these kind of like well-understood jobs that are deterministic that I would say we've got, you know, 90% of the problem solved with Claude Code. But even then, it's still like the case that I think we need to turn it off and act like a commercial airline pilot. Because without the ability to judge the output of the LLMs and really be an expert at it, you know, you lose control of the system. And I think it's always going to be down to us to like make sure that the things are working well.
You say that, and then I have had conversations with aspiring entrepreneurs that are like, "Hey, I saw these YouTube videos and these people have 20 agent different role types and we're going to... I'm going to build a brand new app from scratch and just have Claude build it all." And I'm like, it's hard to argue with that, too. Like well, if they're like... they don't have any customers yet, so there's like a new thing or whatever. Anyways, but I think there's like that: how do you do this responsibly, and you don't know what you don't know.
And then I also remember like 20 plus, 25 years ago or whatever, when I was a new engineer in the thing. I didn't know how to do a lot of things, and the only way I learned is because I had to just overcome whatever new challenges popped up, cuz there wasn't a playbook on how to be a PHP developer or Ruby on Rails developer in that first year or two, when everybody was like, "Hey, you shouldn't be using this new technology." Now we're calling it boring, cuz it's been around forever. But there was a time it was almost seen as a
little dangerous to use something like Ruby on Rails. Like well, how are you going to handle this when the scaling thing pops up or you need a bigger database or whatever and we're like, "Well, I don't know. We'll figure it out when we get there."
So that's what I keep telling everybody that's nervous about adoption of AI is like I think we're all going to figure this stuff out as we encounter it and as long as we know that we're not the only ones doing it alone, if we're going off on some crazy mission on our own and like the rest of the industry is like, "Nah, we're not going to touch that." But that doesn't seem to be the case anymore. I feel like the trajectory very much is like, "Okay, we all need to try to figure out how to get on board or adopt tooling in some capacity."
You're mentioning that is the goal right now still double the PR throughput at this point or is that still the kind of the metric you're kind of leaning on there?
So the good news for us is that we met the goal. Go past it. Yeah, so I think it happened about 2 weeks ago at this point. So we, you know, we roughly said we're going to in a year double the throughput of engineering and/or R&D. And yeah, we met that. And but now we're going to do it again.
You know, the bottlenecks will change. And arguably 2x wasn't even ambitious enough. You know, when you look at where the models are going and the quality of the harnesses and stuff, I think 2x is actually kind of underwhelming. There's a lot of work you've got to do to start all of the organizational change. Cuz like we want to change how hundreds of people work with each other and their craft and their skill. And that's years of work. But we're trying to speed run it.
And you know, the more work we move agentically, the less we're going to be thinking about the kind of lower level stuff. The same way that with compilers, it's like you don't really think about machine code too often. Our work will be focusing on the kind of higher level inputs and outputs. Still have to be able to go deep when you need to. But like why can't it be 10x? This is realistic, I think, to achieve. And yeah, we're going to keep going.
You know, you mentioned earlier that you, with your skills, are using rules as well in Claude or things like that and you distinguish between the differences there. I know a little bit about this stuff with the limited amount of skill development I've done.
Yeah, we use like, yeah, skills, rules, guidance, hooks. And we kind of go through. They're different levers to pull depending on like what you're trying to get out of it. Like hooks are very useful. You can just force a behavior to change, to kick in.
So, one of the earliest skills we wrote was a skill to improve the quality of pull requests, the description. So again, kind of out of the box Claude Code, if you ask it to like describe this pull request, it will do an amazing job, a fantastic job of just regurgitating the information that's in the code. Whereas that's actually useless. And what we found was the quality of our pull request descriptions was going down as more and more people adopted Claude Code because the most important part of the pull request description is the context on why, like what is happening? Why are we doing this thing? And even if you just leave a breadcrumb to like this GitHub issue or something like that, it's like, "Thanks, I know why this happened, you know?" A description of the code in English isn't actually very useful.
And so, we brought in like a rule to call a skill where the skill would look at the session data and go and try and establish the purpose of the code change. And then if it couldn't figure it out, it would ask the person like, "I can't figure out what's the purpose of this."
And that made a reasonable difference. But even then, we found that the things weren't being consistently held and aren't being called. And so, we ended up using a hook to really force Claude Code's behavior at different times. And so, you know, we'll kind of go through these, like see what works, see can we get away with it, forcing all of these hooks or rules or guidance or things. See where we got the things. And if I was make a big difference here, this is like, you know, how you prove it either way whether the behavior happens the way you want.
It can definitely help with that. But then we just use the telemetry, you like ask people if they didn't use the skill, why didn't it? Like are they using the right version of Claude Code? What version of the plugins are on? It's just like a huge kind of tech support kind of thing. But basically involves you just understanding what's happening. Why are people invoking things in certain ways or like why are they not getting stuff? And then understanding the tools and functionality well enough to go, "Okay, we've got this rule in place. It's not being called. Let's move it to a hook." or something like that.
And so we're kind of at the mercy of Anthropic, like whatever functionality, whatever next they're going to release or whatever. But like they're decent enough building blocks to be able to get the behaviors working the way we want. Yeah.
I think one of the things that I find interesting about that is like as you're using Claude Code there, it keeps it a little bit more consistent. But you're also now following that as a moving target. Yeah, like if anyone using even like choosing any type of technology, you're kind of at the mercy a little bit about what they're doing there. And I think, you know, has it been an interesting thing to like, is using these tools allowing you to avoid needing to actually build software within your product? Like say some back-end admin type portal tools that some people are like, "Hey, we need a one-off report thing." or you like, "Hey, can we add this piece of data into the interface so I can pull this thing down?" And you're like, "You only need this number like once every 3 months? Here's a thing you can do from Claude Code." Is that happening at this point?
Yeah, writing code might be the least important thing that's happening to Intercom with Claude Code. The main use case that people outside of R&D have is that kind of exact like analytics kind of, like making sense out of all of the data that's shoved into Snowflake and other tools. And I don't usually like the phrase, but like we truly have democratized access to this stuff in that just in the past this data was stuffed into Snowflake and Tableau and these other places.
And it's just hard to work with, hard to trust, hard to even know, like is this the right fields? Like it sounds like it is, whatever. And you can ask Snowflake, but Snowflake doesn't actually know. So like this, like Claude Code skills for regular day-to-day work that could be an admin tool or something internal. But actually if you go direct with great guidance, you can just pull out the information you want like on the fly, ad hoc. And that's what's transforming huge amounts of work inside of Intercom.
So it's like you don't need all of these dashboards to try and answer your questions. If you could just write your question and then with Claude Code, given the right guardrails and the right access, it just gets it right first time. And so like these dashboards that people would build and have teams looking after them and all this kind of stuff, they're kind of obsolete almost in a certain way now because like what is the job they're doing? It's like people have questions about things. They just want to get the answer. Now they just ask the question and get the answer, which is, and this just seems way bigger than, yeah, we can produce like database migrations reliably, you know.
Yeah, that's interesting cuz I think about just the number of different client projects that I've worked on over the years where we're building out little data interfaces or reporting dashboards or something. And they're trying to figure out how to optimize those interfaces even though they don't always log in to even access those, but they need that data for their weekly leadership meetings. They can plug it into some KPI spreadsheet or something. And now we're like, "Can we just find you a way to get that data some other way?" Or can we send it to your sheet? Or can you fetch it through like a written out prompt or something? It's kind of like you're providing different sets of endpoints into your application and it's just maybe not through a traditional API. So maybe kind of moving back into the Rails ecosystem, you mentioned Snowflake a couple times. Is that like integrating with like the Snowflake CLI tool that it relies on and then there's just some API keys, or providing things through your Rails app that then it's fetching and kind of playing a proxy, or is it a bit of a combination of a couple of those different types of things?
Today mostly we have standalone MCPs. Like there might be some CLIs out there but we use just like a local MCP or individual MCPs for different services. So, you know, means I have to auth to a million different things very often to do my job.
But we've also been doing a bunch of consolidation into our Rails app as well. So we've had these admin tools for as long as Intercom has existed. Interesting thing is like when I joined Intercom I saw like our customer support team, say, using the Rails console and User looked up. Yeah, they would just like, it was like the way that they were querying production. They would like, you know, just go in there, type in some codes, type in some stuff in the REPL. All good.
But then as they became more specialized and we became bigger and became more boring, you know, we locked people out of the Rails console. We obviously matured and made it kind of a safer place to do stuff but it started being somewhere that only engineers kind of went. Look, there's lots of good reasons for this and if you can write a quick web page to do something it's probably a better way of getting some information.
So one of the philosophies we have is that anything you can do on your laptop, your agent should be able to do as well. And one day I was just doing some queries on the Rails console and I was asking Claude some stuff about it and then I was like pasting the output into the shell. I was like, what am I doing here? Like why am I this proxy for Claude Code to be able to run code in our Rails console?
And so we were already starting to build API versions of a bunch of internal tools to do, say, customer lookups and pretty well understood normal things. But I was like, what if we let Claude Code actually just log into the Rails console and do it over API? Like do it over MCP. You can send arbitrary code, like building on top of all of the kind of safeguards we have in place already in the Rails console.
And so I built it and even I was a little apprehensive, like this feels a bit weird, and kind of quietly launched it and then started looking at usage and the top five people who are using this thing, they're running like hundreds of queries. They were just doing ad hoc exploration. They might be pulling in data from Snowflake, pulling in data from different places. And every so often they would like want to do something and Claude would figure, you know, the best place I can do this is on the Rails console and I can just call this API to like do this little thing, do this little transformation, write a little bit of code and get the exact answer I want rather than like maybe trying to rehydrate it or like sucking 10 different bits of data from Snowflake or whatever. And the top five users were all non-engineers. I felt so proud.
It's like these folks would not have gone on to the Rails console to do this and they're kind of just prompting Claude and Claude is figuring out how to get work done for them. As a result they're just doing a bunch of stuff on the Rails console via Claude Code. And so it justified, I guess, the approach that we had, which was like, again, anything you can do your agent should be able to do, and then be like maximizing that.
So don't just allow certain people to access Snowflake or don't only allow, say, certain people to access AWS. I mean I'm not saying be reckless here but there's no harm in having a read-only view into the entirety of your production environment. There's very little that can go wrong as a result of that. And so why shouldn't everyone have the ability to get Claude going with a read-only view into production. Similar with the Rails console. If you've got like a safe environment to run arbitrary code, why can't you let customer support or like, I've had design managers and senior directors of product unwittingly running fairly complex Rails REPL pieces of code just because Claude Code is kind of doing it for them, and it buys so much productivity and it's even hard to know that it's there. Like if I asked our head of product who was in there accidentally using the Rails console, I was like, hey, do you need access to production to answer these questions? They probably wouldn't even know to be able to tell you. But Claude can do it for them.
And so being like bullish or being very liberal in access to things and setting up things well, like you can really give people so much productivity and improvements in their quality of life. But you've got to be able to push through like the apprehension around giving access to production and that kind of thing.
So but in that workflow you did add some observability into that workflow and like so you know when people are doing this, so you have to stay compliant. You mentioned like SOC, things like that earlier, like who has access to the information, who accessed the information, when was it accessed by, you know. I could imagine some people might be like, well, that sounds easy enough to say but like at the end of the day someone's computer gets hacked, you know, and like you just gave access to all your production data. What's your kind of counter to that? Is that just like, well, that's just kind of the risk of doing business with anything that gets installed on someone's computer? What is that kind of compliance concern? How would you answer that question to someone listening right now?
Yeah, so for these cases, like these people had access to all of this data in the first place, like prior, and the Claude Code is kind of acting as a proxy. And for sure it accelerates the volume, the amount of times that they are accessing us, and so gives more entry points where things could go wrong. From a strict compliance or risk management point of view, doesn't change much in terms of like what actual data people have access to.
So your job is still the same, you know, detection becomes more important because you're expecting more people to be kind of practicing this, you know, people's laptops now have more credentials, more active sessions into things. Like, because they're being more productive, so there's more surface area, I guess. And then there's the whole class of issues with like securing at an arms in general.
Yeah. It's very much an unsolved problem and if you control the inputs or you control the harness or whatever, like you can control the outputs. There's definitely like new risk that could be involved as well and this is where governance, you know, making sure which MCPs people can connect to, where people are getting their plugins from, reducing the risk of like kind of garbage going in and losing control of the system that way. So we've been doing work there.
And, you know, it involves having robust setup in terms of detection on people's laptops. We distribute plugins and config through our internal IT systems. So we'll like lock down Claude in certain ways and make sure it can't do certain things. Does this stop somebody from downloading Codex and bypassing all
this completely or whatever? Not at all, but it makes the default kind of safer and stuff like that. So I do worry about this. It's like the pressure on the CI/CD system. It's like you're getting this pressure due to amazing success. You know, people are getting more work done. That is great to get. You just want to make sure that the kind of sharp parts of it are well mitigated and, yeah, we just got to be constantly increasing the amount of vigilance that we have in terms of like the broad security system and compliance control system because the system's doing more work.
You know, some of the latest news that's been coming out of Anthropic and others about like new models and new security threats, things are going to get more exciting for security teams over the next while. I think it comes down to doing the fundamentals well and, you know, you don't want to encourage shadow IT. You want to make your defaults set up great and give people access and control, but also give you kind of confidence and auditing and all that kind of good stuff that you need to like make sure you're not being completely reckless or abandoning
You know, I hadn't thought about this aspect of like having people that have historically been software engineers and like maybe opening up a terminal for the first time to open up, you know, Claude Code from the CLI or something, and are you finding that the adoption there, like people are kind of generally, are they intimidated by that type of UI or the terminal user interface, or is it kind of like it's a little exciting to them, and what's that kind of look like over there?
Yeah, you know, the funny thing about this is that like how I started getting into computers properly like 27, 28 years ago was I used to run Unix systems that were run by students in the university I went to. And so we had newsgroups and IRC and instant messaging and email with our hosting. All that kind of stuff that you'd do in like 1998 in my social time. And it was great. We were the most successful society on campus. The university I went to was very like internet friendly. Everyone had internet access, which wasn't always the case. What we had was like the best social network. And it was a great way to meet people. It's a great way to troll people, you know, it's all the stuff that happens on the internet. And so, we had like thousands of like non-technical students. Like, they would be banging down our door looking to learn how to like use the terminal so that they could do all of this cool stuff that they'd see their friends doing.
And now like nearly 30 years later, it's like my job is exactly the same. I'm teaching people how to use the terminal because they see their friends or colleagues use them and go like, "Oh my god, I need to like get started here. I don't want to be behind the times."
But, we're also seeing loads of like really interesting ground-up stuff. Like, our VP of operations recently released like a guide to getting started with the terminal. Like, he sat down with Claude Code and like I think he had been teaching a few people that he worked with how to do this stuff. So, he just like built this little web page or, I don't know, a little application for other non-technical people to get started with Claude Code. And, you know, he described it as like, "Look, this terminal stuff, it's real nerd stuff. Like, and it's very off-putting. But, you know, just use my guides and you'll be great."
So we're seeing this like amazing bottom-up enthusiasm where if people have the motivation, and I definitely saw this in college and I'm seeing this now, like the terminal isn't that off-putting. People can get used to it. They just need to know that they're going to get something out of it. And if people are seeing the kind of productivity wins, they'll figure it out. They'll like get in there. Now, I don't know if they're going to be there like, you know, installing Oh My Zsh, like mastering the terminal. We might get one or two, but like you motivate someone or you give somebody motivation, they'll learn it pretty quickly, you know.
Thanks for reminding me of like the late '90s era, cuz like that's when I got into this stuff as well, and like it was like, "Oh, I want to connect to IRC channels or I'm telnetting to some thing, you know, or connecting to some BBS." And getting used to working in Linux machines, and then fast forward several years, you know, I think the era of when I saw a lot of adoption of terminals again was getting web designers and developers, like front-end people, comfortable with the idea of using a terminal so they could use the GitHub CLI, you know, or just using the GitHub CLI tool. And that's when Oh My Zsh, you know, came about, was cuz I just wanted those people on my team to feel more comfortable interacting with the terminal's interface. Like, this is going to be so much better than this GUI GitHub thing that you could be using. And now there's this whole new generation of people coming in, of like, terminals look and feel a lot more comfortable now. And there's a lot of amazing stuff there, but underneath the hood there's still like a bunch of CLI tools. And so, you mentioned
Anyhow, I was just kind of curious also, coming back to the Rails part with the CLI. How does that connection work then? And you mentioned like there's some MCP servers and stuff like that, but how does Claude know how to talk into the Rails? Is it a remote console, like almost like a, you know, something really simple thing that maybe a lot of Rails developers, like, you can do like a Heroku run console and Claude could in theory feed some stuff into that. What does that look like? How is that working, and is it connected to your Rails production environment to be able to do things?
I mean, I didn't look at the code. I just told Claude to do it. No, I have read this code though. It's pretty sensitive. You know, it is effectively taking the string that is injected as an API parameter, and it can take a multi-component string. I don't know what the exact terminology is, but you know, it can have all these compounding bits that are called kind of independently. So, it can take effectively a function or whatever. Like almost like a lambda. And it runs a bunch of checks against those to make sure it doesn't do this, that, and the other. And then it attempts to run it, reports the outputs, logs it in the audit trail. So, yeah, effectively like whatever valid Ruby code comes in, it'll do its best and try to just execute
Does that require that person to have the repository, like your main Rails app, cloned locally, or are they able to just open up Claude from their home directory and then like have access to your skills and talk directly?
Yeah, now this is running in our production environment. So, they're connecting the same way as if they were SSHing into one of our hosts and running like, you know, a Rails console. It is the same level of functionality. Yeah, and able to take in just arbitrary bits of code. It's not as powerful as like the console REPL. You know, you can't tab complete to find all sorts of things. It's probably going to run on a different server every single time or whatever. So, you kind of got to get it right first time. So it's not as good an experience, but Claude can call it. And that's where the power is. And like, you know, it gets it wrong many times. It'll send all sorts of invalid stuff over. But, like when it gets things right, it's amazing to see.
Interesting. So, like someone that's not in your Rails engineering team, they can fire up Claude and then they're looking at that locally, that is connecting to the production environment console, and then it can do some stuff there. How does the local context of Claude know about the code, or is it running the whole thing? Like, how does it know to like user.where whatever, like to generate the appropriate Ruby code if it needs to generate something new, or is it like running specifically a specific method that's already been coded?
So we give hints. So, we have like a skill or guidance that goes along with the call. And so, that'll give us some examples, you know. But then generally when people run this, they're in the Rails code base. Claude from there could just explore the code base, figure it out. And that's exactly how I've used it. You know, I'll be like asking some questions. Say, "Hey, I'm looking for this tag, I want to find all the places where this tag is." And so it'll be kind of back and forth between like explore the code base, run some stuff in production, explore the code base, run some stuff in production. So it's not just the running of the Rails console commands that it'll be doing. It'll typically be done with either looking at other information or, yeah, exactly like the code base that's there. And like that works generally well. Like, it doesn't get stuff right first time. And there can be like gaps where it's like going off and exploring the code base and stuff. But it gets there. Gets there pretty fast.
I don't remember who I was talking to about this, but there was kind of like this idea that, say, let's fast forward a year or two. Might there be scenarios where you want to ask a question of your data set in your Rails app, and normally you might build some code to reproduce that, and might there be a scenario where Claude or something will temporarily create some classes and methods in a Ruby thing, run it, overlay it like on your Rails app, and then answer the question, and then that code just goes away because it's never needed again. And then like, what's the purpose of keeping code around if you don't need it? Like, how do you remove some of that dead weight in your code base to keep the context window smaller? Do you think something like that could happen? Do you see any advantage to that, or do you feel like it's just too pie in the sky?
No, I think it's interesting. Like, if you look at our rake tasks area in our monolith, it's like ephemeral stuff typically. It's like a lot of one-off stuff, but you have to put it there somewhere so that the code gets into the right place that you can run it. Getting back to code reviews, it's a good place to like give feedback. Or if people want to like check that their task works, they can get a review or whatever. Yeah, it's an interesting idea that we would have more transient code that may not even end up into like Git or whatever. If it truly is a one-off thing, then like, well, maybe this Rails console MCP might be where it ends up going and that's what it does.
So, like, you know, you always want an audit trail. You definitely need to see what has run. Like as an execution environment, this sounds kind of appropriate for like these kind of ad hoc tasks or whatever. I probably want to invest more, make it a bit more stateful, make it a bit easier to get right or something like that. Maybe more helpers. Also, we've got it pretty locked down. Like, you can't mutate something in the Rails console using this thing at the moment. And we're not like morally opposed to it. I think we just want to enable this stuff when it makes sense. And changing to like allowing full writes, we just haven't seen the need for it yet. Like, it hasn't been asked for. And so, we're very happy to kind of give read-only access as such. And then writes, we're kind of still happy to... Like, if it's important enough to have something you're consistently writing on, then build an API endpoint. It's like one Claude Code session. And you can have all the validation you want and stuff like that at that point. You get something that's that bit more predictable.
I think it would be amiss not to ask, you know, you mentioned your team's kind of at least made the decision like we're going to work on Claude Code for now. That problem is solved for now. Let's just get as far as we can with that, and I'm sure you would pivot if necessary at some point. But, it would have to make sense for that. But how is your team thinking about token consumption and projecting potential future costs? Cuz some people are like, "Is it premature to try to be thinking about optimizing this stuff right now?" How are you thinking, like, what's the current status of... And I realize by the time this gets published your opinion may change on this. So, middle of April, where are you at, Brian?
Yeah, we think the opportunity costs of caring about optimizing token consumption is too high right now. So, we are not thinking about token consumption. Now, that said, it's an extraordinary amount of money that we are paying in tokens. You know, look at our traffic's revenue. And I think it's appropriate for us. We are very aware of this, though. Like I think the growth rate that we're at right now is not sustainable. But, opportunity cost where we're at and wanting to be on the bleeding edge is worth it for this band. And we have like some pretty short-term wins or sort of things that we know we could turn around. But right now we're just not going after them.
I also think the harnesses are getting better, even just some features this week that Anthropic released. They're very aware that people are spending a lot of time worrying about tokens. And it looks like they're getting smarter as to which model to choose here and there. That's good. I'm sure there'll be more improvements in terms of caching and different things. Like I think our token spend will just continue to increase and have no sign of ending there.
But, the main thing that I want to make sure is that we're not being sloppy, that we're getting value for the tokens. That it's not like unnecessary burn. We're not like unnecessarily using overly complex orchestration harnesses that just burn tokens for fun, or that we're like pointlessly using Opus when it's a perfectly good thing that Sonnet could do or something like that. Yeah, I'm kind of at this point just making sure we're not being caught out by obvious kind of things. I think that we've seen Playwright in particular as well just like burning through a bunch of tokens. And with some minor changes, you can make some good improvements there. So, I think we will be spending more time over the next while on like improving this stuff. But today I think it's a very compelling answer to just go, "Look, just spend some time worrying about this in a few months' time."
But, Brian, like I got to answer to my finance people. They want to know how much this is going to cost. And like are you saying this is just going to keep getting more expensive? How can I manage my budget? Like are we talking... I don't know, I think I say that kind of somewhat specious, specious, I can't speak the word. But you get the idea. But it is like a real thing that some people are challenged by, like if they're not a company that's historically been on the bleeding edge of technology and see that as part of their competitive advantage. And they're like, "Oh, we're trying to figure out how do we adopt this tooling?" For those listening, do you have any advice on how to navigate the conversation so it doesn't seem so... I mean, outside of being like, "Yeah, upgrade to the $100 or $200 a month Max account." Like if you can't even convince your finance people to give you that budget, like you're probably... But it's also not just that cost, as we talked about earlier. It's also the cost of the infrastructure, your CI pipeline, and increased cost as well. So, and it might be difficult to project those estimates if you're kind of a small team. Yeah, like
In many ways working in Intercom is like doing this on easy mode because, you know, we're completely all in on AI. Our CEO, like all our leaders, we all completely get it. And our finance team are like big users of Claude Code at this point. Like we've had a lot fewer questions since they all started getting Claude Code-pilled.
I do a lot of work with finance, our finance team, through, like, back in my old life I used to take care of our AWS account, which is actually our biggest source of spend in Intercom. And in one of my regular chats with our finance folks, I gave them a Steve Yegge article. And I think it was called Death of the Junior Developer. And he was predicting, as Steve does, he had a few predictions. And one of them was like token spend. And he was saying like, "Hey, you know, we're at this point, like it's going to be 10, 100x in the next year or two." And I was like, "Damn, this Steve Yegge guy is pretty good at predicting this stuff." And so I sent it on to my finance guys, cuz they were asking like, "When's the growth going to end?" And this was like a year ago. And I'm like, "This is only going to get worse and like a lot worse." So, I kind of feel like I primed them well enough.
And yeah, we got them Claude Code-pilled, which helped. But, more fundamentally or more importantly, we kind of think of this as this should come out of your headcount budget. This is not like sunk costs or like, you know, treating it like some seats to JetBrains or whatever. It's just like a more profound kind of tool. I mean, you need a different way of thinking about how to pay for this stuff. And I think like raw headcount is a good way to put it. Also at the same scale of it as well.
But then the other thing is your company really have to understand how transformative and what a big deal this is. And it's your job to really break through the barriers, the controls, the concerns, all of the things that kind of prevent adoption. And one of it is like justifying token spend and stuff. And look, I think we will do some work to improve token spend and get more control over it. So, there's loads of things that might end up happening. But like today, I think most businesses, the vast majority of businesses, and certainly any kind of progressive tech business just has to be running as fast as possible to this stuff.
And if your finance team or whoever's in charge of the budgets aren't doing the right thing here, that needs to be immediate escalation to C-level or above or something because it is so critical for us all. Like you know, I wish us all the best of luck with all of the change that's kind of going through this. And like you either have a choice to do this stuff now or do it down the line when you may have already missed the boat. And so I think it's correct to act urgently with this stuff. If it is old school bureaucracy and budgets and stuff that are getting in your way, then you need to move as fast as possible to blow those things up because you're going to look pretty bad in a year or two if that's what slows you down or made you roll out more conservatively at this time.
How do you feel like Rails now fits into that? It feels like that feels maybe a little at odds with, as you earlier described yourselves, technically conservative, yet innovation, you're trying to stay bleeding edge in these realms. But, do you feel like Rails is providing you a good foundation to be able to, like, we got this boring, I don't want to call Rails boring necessarily, but this boring stack and the other technologies that you rely on there at Intercom, you have that kind of settled. Settled, you're not looking to transition away from that. That's my assumption anyways. Now we can take some steps up on top of this and stand on that work.
Do you feel like Rails and Ruby has proven to be quite useful in this? Do you feel like you're really, really fortunate that you made the decision as an organization to use Ruby and Ruby on Rails early on, given where the technology is? Cuz I think if I go back like a year or two ago, something that I had heard was like people were like, "Well, if the LLMs are trained off of available software, you know, is it going to be suggesting that you use other technologies because those technologies are the ones that are most available?" And Ruby on Rails apps are not necessarily the most popular, there's not as much of an abundance. But, the context window is a little bit different when you're just looking at the code itself. So, that's a long convoluted question, I suppose. But, tell me how you think LLMs and Ruby maybe do or do not complement each other.
Yeah, like if you asked me two years ago or so, I definitely would have said that we see better results with say writing React or Node or Python compared to writing Ember. We have a lot of Ember code. Or Rails. And especially kind of custom Rails, or the kind of Rails that you end up writing in a big code base, just something seemed a little off. The harnesses weren't kind of grokking things correctly. Claude Code seems to have just vastly improved on that. So, I think like the discovery mechanism, the way it kind of, you know, just grabs your code a lot. Just kind of picks things up and it's very fast on that. So, something did change last year. That didn't seem to be only the models. The harnesses just seem to be getting better at dealing with larger code bases. And so then we saw that as it's just starting to get a bunch of Rails stuff in our environment right first time.
You know, I saw some recent study that did kind of make a case for Ruby being particularly friendly for LLMs. I think if stuff is good for humans, it's good for LLMs. I think I have a simple enough approach to these things. If you write docs, it's good for humans. If you write docs, it's good for LLMs. Same with tests, same with like short pieces of code. You know, I think all of these things, there's many similarities to like if you do stuff that's good for humans, the LLMs will have a good time, too. And I think the kind of craft, the design, the way you can get stuff done pretty fast in Ruby and Rails, the expressiveness of it, I think all those are well-geared for LLMs to be able to just write good code.
Maybe it was like the LLMs nailed like Java boilerplate style applications or something in the past because it needs to be more verbose and they seem to be well-suited for that, but I think the harnesses now just consume to run against many different styles of languages, and I don't see it getting things wrong. I don't see it going down completely wrong ways in our Rails codebase, and I've just had some good experiences of asking it, "Hey, take this code, make it a bit more idiomatic." And you know, it just kind of does it. It's not getting stuff wrong. It's not like you're seeing a Python coder or a Java coder try and write Ruby, and so you look at it and you go, "Ah, it's not bad, but not quite right." When you point it in the right direction and you give good context and tools and all that kind of stuff, we're not seeing any problems with quality of code or completely wrong approaches kind of happening in our codebase.
So, I think the future is bright. I think maybe in the early days what we saw were maybe signs that it being an obscure-ish language was holding things back, but yeah, I think in the last year or so we've just seen things break through, and I haven't used many different Rails codebases. Maybe it's different if you've got something structured unusually. Maybe even if you're using Sorbet or whatever, maybe that might be a bit of a barrier, but yeah, Rails can navigate, sorry, Claude can easily navigate our codebase and produce more code, figure things out pretty fast, write really good idiomatic stuff. I think the teething pains we had maybe 2 years ago, I think we're long past that now.
Before we had hopped on and hit the record button, you had mentioned like, "Oh, it'd be nice to be able to advocate for some things that I would love to see Ruby and Rails provide the community to help us with some of these things." So, let's fast-forward 6-12 months. What would be some of the ideal things that you would love to see kind of emerge in the community to help provide tooling like this for other organizations and for Intercom?
Yeah, I think once you start considering the agents as the primary way that is deploying the application or picking the technology or working with us, like agents to app, kind of figuring out telemetry, observability, customer experience, whatever. This kind of covers a few things that we touched on already, like the execution of ephemeral code, giving the ability for agents to answer their own questions or figure out how to get started or where to go. I think thinking of what an agent-first experience should be for the Rails ecosystem means kind of starting from first principles, like how do people decide to even install an app or install a framework.
It's not hard to get started with Rails. I'm pretty sure if you tell any kind of modern LLM or whatever, "Hey, write me a Rails app." But why isn't Claude picking Rails first? I was listening to the Cheeky Pint podcast the other day, and I think it was the interview with Google's CEO, and I think John Collison mentioned that he'd been working on some side project, and then, you know, he was kind of finished, but then he was like, "What language is this?" He hadn't even thought about it until some point down the line. And I was like, why doesn't my LLM automatically just pick Rails as the default and stuff?
So, I think that's kind of an interesting avenue to go down, as like in the installation and setup and documentation and all that, that kind of zero to one thing. Thinking of that as an agent-first experience, what can we do to make that just work out of the box? Cuz I think at the moment it picks Node or it picks, you know, React or whatever. These are just the things that they kind of default to. And so, it'd be interesting to see if we could contribute to optimizing for agents so that it's just picking Rails as a more natural thing, and it kind of happens organically or whatever. So, agent SEO, but also like agent experience or something like that, I think it's interesting. And we've been doing a bit of work on Intercom to kind of make Intercom interesting to agents as well in the same way.
Interesting. Agents as the interface. Like webpages, dashboards, all these things, they're very interesting for people when they're the end users, but an agent, same way, all the things that you would expect, like the Heroku CLI, to be able to interact with. And it's going to be other agents asking other agents what's going on, what errors have we seen today or what usage or whatever, and then taking action on those kind of things. And so, you know, I would love not to have to write these APIs for console API. So, having these kind of things either native or well-supported community gems that add in like Console1984 gem safety, and defaults to really good patterns to make it so that if you do choose Rails, or maybe your agent chooses Rails, it's the best agent-first environment for you to use after that point, and the agents can do everything.
And the kind of smarts that we're putting in place, like convention over configuration in Rails, I think some of those could just be reassessed with like, "Okay, let's assume it's not people driving this stuff. It's going to be agents driving them." Then what are the tweaks or what are the kind of optimizations? Or maybe we even have to throw away some of the conventions to be more agent-friendly. So, it's profound. If you just assume that it won't be people doing Google searches or looking at Stack Overflow or whatever trying to figure out what to install, it's just going to be agents doing all that. And so meeting them where they're at and running stuff down the line as well. So, meeting them where they're at and making it compelling because, you know, the agents got less, and they will do whatever it takes to solve the problem being given to them. And if the path of least resistance is to configure a Rails app or to turn on something that's already configured in Rails rather than doing something else, you know, then that'll make it more compelling to kind of use them. So, I think agent experience really matters, and it's not just true of SaaS companies or frameworks, like all software vendors, everyone needs to be kind of thinking about this.
How do we solicit feedback from the agents to make sure that we're giving them a good user experience?
It's a good point. You know, if somebody asked me that internally in Intercom, I'd be like, "Oh, just write some evals." You know, it's like figure it out. Write some tests, and then figure out where the breaking points are or something like that. And certainly that can kind of give you a benchmark over time, I guess. You can ask them sometimes. You can get some useful stuff out of that, like dive into your session data and answer why you kind of decided. But honestly, I don't think they're, they're just not self-aware. They'll regurgitate a plausible answer, but it may not actually describe what's going on. I think some of it's like basic SEO type stuff. But also there's stuff that we can do, like making developer docs accessible over markdown, making search accessible, things like that. So that it's not just you end up on a cool-looking landing site, it's like your agents need the same experience as well, that when they go and look for some docs or whatever, it's super fast, and they just get there in fewer steps.
I didn't, yeah, think about that at that lower level there of, like, if anyone has an API right now and they provide API clients, but also provide, like, you know, DHH is talking about how they're building some CLI tools, and Intercom I think is doing the same, and everybody's starting to build some CLI tools for their APIs. How do you optimize those doc sites? And it's like maybe you don't render them the same way because you don't need to have all the visual stuff. You want to make that as quick as possible because it takes more tokens, I would imagine, to read a webpage that's got all these fancy graphics and little visuals and videos. It's probably not going to sit down and watch a video.
No. Yeah, but you can summarize and give a transcript and all this, and it's kind of like latency or something like that. You know, people talk about how many people drop out of a shopping pipeline if there's latency here or whatever. It's similar with agents. It's like the more tokens you have to burn, the more steps. Even if they get there in the end, if you can reduce that down to as few steps as possible, you know, you'll get a way higher conversion rate
equivalent or something like that. And I don't kind of tell you right now like that will obviously mean that will influence the agents to do it, but like certainly if you don't, I think it might cause to like exclude whatever. And I think there's a good case to be made for like Ruby on Rails being very LLM and agent friendly. So, we should do the rest of the stuff to make sure that they can adopt it really good.
Yeah, appreciate that. All right, Brian. I again kept you far along enough, but a couple of quick last questions for you. Is there a non-programming book that you'd like to recommend to peers to check out that feels appropriate in this time and place in the world?
Yeah, I'm going to go with a technical book. And I mentioned earlier on just like a few blog posts like Dan McKinley and stuff like that. I like re-reading the classics. By classics I don't mean like actual classic literature. If you just know a bunch of core material pretty well and you keep it kind of top of mind and you remind yourself the existence of the stuff, you can go a long way in tech. Like you don't need to read every book. You just need to read a handful of the good ones.
And I have to admit getting out Designing Data-Intensive Applications recently enough and just having a quick pass through it. I think there's an update out or something like that. I think that might have prompted me to getting it out. But yeah, I don't know what I learned, but I learned that I need to make sure I recall this book and use this in my day-to-day work and just like never forget the fundamentals. So, yeah, sorry, bit of a cheat answer there going for an actual technical book. And it's a book I've read many many times, but like I think it's interesting to like make sure you're still fresh on the classics.
All right. One more time with the title and the author for that one?
It is called Designing Data-Intensive Applications and the author is Martin Kleppmann.
Okay, great. We'll definitely include links to that in the show notes for our listeners. And with that, Brian, thank you so much for swinging by to talk shop with us on On Rails. It's been great fun.
Thanks so much for having me. It's been a converted chatting with you again. Hopefully we'll see each other at a conference again soon.
Totally. All right. Cheers.
That's it for this episode of On Rails. This podcast is produced by the Rails Foundation with support from its core and contributing members. If you enjoyed the ride, leave a quick review on Apple Podcasts, Spotify, or YouTube. It helps more folks find the show. Again, I'm Robby Russell. Thanks for riding along. See you next time.
Article published
