How Ramp Moved From Many AI Agents to One Agent With Many Skills

Open on YouTube ↗
Overview

At The Pragmatic Summit, four Ramp engineering leaders described what the company learned building AI products for finance teams: Nik Koblov (EVP Engineering), Veeral Patel (Director, Applied AI and Spend), Will Koh (Staff Engineer, Applied AI), and Ian Tracey (Staff Software Engineer). The session covered four things:

24 min read
  • A strategic pivot away from building many separate agents.
  • A detailed account of how Ramp built its policy agent.
  • The internal infrastructure behind Ramp's applied AI work.
  • An argument about how engineering culture has to change when coding agents take over more routine work.

Koblov framed the talk around what the speakers see as a paradigm shift in software that "requires complete rethink," and, with it, a simplification of the stack.

The Coffee Problem: Where Ramp Started With AI

Koblov introduced Ramp as a finance platform for modern businesses with more than 50,000 customers, whose business is "saving you time and money." To illustrate, he used a cup of coffee bought on a company card. By his account, the purchase usually costs an employee about 15 minutes, because of three simple steps that each take minutes. That cost compounds across a company.

Ramp handles that whole chain agentically, Koblov said:

  • writing a memo after the card is tapped,
  • classifying the transaction against the general ledger,
  • sourcing and attaching the receipt,
  • normalizing the merchant against the company's merchant inventory.

This was Ramp's first use of AI, starting roughly three years ago with "one-shot" tasks such as normalizing a merchant or writing a memo. Koblov said these have worked "really, really well as the models get better."

From there, the ambition grew. Koblov argued that nearly every finance-related role is losing time to manual work, including AP clerks, finance teams, purchasing teams and data teams. The complexity "has a ramp shape," increasing as you move through different jobs to be done. His example was a Slack channel called "help data," where someone would ask for a CSV and "a poor person will go and write a SQL query." Ramp replaced that process about a year and a half ago. The company's end-state goal is to cover everything admins, employees and finance teams do that is not directly related to making money.

The Pivot: One Agent, a Thousand Skills

The first major lesson Koblov described was organizational as much as technical. Last year, Ramp intentionally let each team experiment on its own. The result was "maybe four different ways of doing the same thing," for both synchronous agents and background agents. His conclusion: "you don't need to build a thousand agents." Instead, a company should drive its framework toward "a single agent with a thousand skills."

To explain why, Koblov broke any modern AI-driven process into five parts:

  1. an event, such as receiving an invoice that needs paying;
  2. prompt instructions describing what to do, plus guardrails such as an expense or payables policy;
  3. the context the agent should consider;
  4. tools;
  5. the APIs and actions that can be taken.

In his framing, traditional software focused only on the last two, the tools and actions. In the new paradigm, "software is doing everything." The goal becomes an autonomous system that can react, reason and act with little or no human supervision.

In practice, the first step was consolidating conversational interfaces. At the end of last year, Koblov said, Ramp had about five different conversational UXs. These have been merged into one, called OmniChat ("Omni" meaning omnipresent), which is being deployed to every surface of the product. He stressed that it coexists with traditional UI, because "you still need tables and buttons" and people don't always want to talk to their software.

His demo was a request to onboard a new employee. OmniChat resolved the person to an employee ID and looked up their place in the corporate structure through an HRIS tool. It then found a previously created agentic workflow called the "new hire playbook" and asked whether to use it.

This runs on an in-house, lightweight agent framework that orchestrates tools. Koblov said engineers build these tools very quickly. More recently, one product manager vibe-coded about 20 tools, so engineers are "no longer needed to build these tools."

Playbooks handle more involved workflows. For onboarding, a user can describe in plain language what should happen when someone joins:

  • give them a card,
  • make sure they submit receipts for every transaction,
  • congratulate them on Slack,
  • check in after two weeks.

Ramp compiles that description into a runnable, deterministic workflow and hands it to the agent to execute.

Koblov then showed how the pieces connect when a card is swiped:

  1. A real-time policy review runs, enforcing the company's spend requirements. He said this is what makes it safe to give cards to every employee.
  2. A handoff goes to an accounting-code agent that classifies the transaction using the finance team's rules. Traditional products would expose that choice to employees, who usually don't know how a transaction maps to the GL. The agent does it better, Koblov argued, because it has the full chart of accounts and understands the ERP.
  3. The agent either auto-approves or, in the worst case, brings in a human to review for materiality or to flag out-of-policy spend.

What the Policy Agent Does

Veeral Patel took over to go deeper on the policy agent. Finance teams look at receipts every day, sometimes hundreds or thousands of them, and Patel admitted he would likely make mistakes judging them by eye. He walked through three examples of the agent's decisions:

  • A dinner receipt. The agent reasoned over the image and the transaction data and found eight guests on the receipt, which Patel said he could barely see himself. It confirmed the cost was under Ramp's internal $80-per-person cap and that the occasion was a team welcome dinner. With the amount and merchant verified, it recommended approval.
  • An OpenAI charge. A colleague had been testing ChatGPT features. The agent judged this a valid business expense and recommended approval.
  • A $3 bakery charge. The agent rejected it because it wasn't part of an overtime purchase and didn't happen on a weekend.

The product grew out of a customer request. A Fortune 500 customer asked Ramp to approve certain types of expenses and reject others, and brought a list of rules. Patel had worked on early versions of Ramp's deterministic rules. He said the team chose not to keep adding incremental rules. Instead, they took a page from Andrej Karpathy's line that English is the new programming language and turned the expense policy document itself into the rules. He showed Ramp's own expense policy next to a production screenshot and said the product is seeing "really great" use.

Building It Like an Early-Stage Startup

Patel described running the project like an early-stage startup inside Ramp. The team found design partners, including that Fortune 500 company, and held weekly meetings with each of them to collect feedback and decide what to improve.

The main cultural lesson, in Patel's view, was accepting that "AI products cannot be one-shotted." Everyone on the team (PMs, designers, engineers) had to agree that the product would not be perfect on day one.

The team dogfooded internally and began with an even narrower problem: whether a "coffee with a colleague" expense should be approved. Ramp's finance team considered these low-risk, single-dollar-amount transactions.

An early finding once the agent reached production was that wrong decisions stemmed less from the models and more from the context the models were given. Patel said the team could have tried to list all the needed context before starting engineering. They decided it was better to learn from live internal data. One example: an employee's role and title matter a lot when reading a policy, because C-suite employees may have higher limits or be allowed first-class flights on certain routes. So the team began extracting more information from receipts and pulling in HRIS fields already available in Ramp.

From a Simple Pipeline to an Agentic Black Box

Will Koh described how the architecture evolved. The team initially "went big," aiming to automate all of finance and all reviews. In practice, they started small with questions like whether a cup of coffee fits the expense policy. Koh explained that even a simple-sounding question grows complex. Even if the team had gotten the context design right up front, he expected it would still break once generalized to other businesses. His principle: the simpler the system, the easier it is to iterate on, and complexity can be layered on once you know what works.

The first version was a classic pipeline. An expense comes in, the system retrieves context, and a series of well-defined LLM calls asks whether the expense is in policy, why, and how to show that to the user. The output is then formatted for the user.

The next stage recognized that expenses differ. The system classified each one (travel, meal, entertainment), used conditional prompting, retrieved context based on the category, and gave the model some tools. That let it decide on its own that it needed, for example, flight information or the employee's level.

A few iterations later, the result was a full agentic workflow. The agent has complex read tools across the platform, drawn from a company-wide internal toolbox that all of Ramp's agents share. It also got write capabilities: it now records decisions, writes its reasoning, and auto-approves expenses on users' behalf, all in a loop.

Koh was explicit about the trade-off. As a system goes from simple to complex, capability and autonomy go up and the AI "seems smarter," but traceability and explainability go down. The team can read the model's reasoning tokens, but "in the end, we have no control over it." A small black box becomes a bigger one.

His response was to build strong auditability from the start. Assume you know only the inputs and outputs, and ask whether you can verify the agent did the right thing from those alone, even as the black box changes.

Users Aren't Ground Truth

Koh said Ramp initially assumed, as with many products, that users' own decisions would be correct: if a user approved, the agent should approve. It turned out users are often wrong. They may not know the policy, they may trust their employees, they may be lazy, or "it's a Sunday." When the agent copied user behavior, finance teams would come back and say an expense shouldn't have been on the company card.

So Ramp had to define its own standard of correctness. The team held a weekly cross-functional labeling session, which Koh said produced two benefits:

  • A ground-truth dataset the team trusted and could always test against.
  • Shared understanding. When the agent got something wrong or was missing context, everyone knew. That reduced communication overhead and aligned priorities.

The sessions were expensive, though. Getting people in a room weekly and assigning homework of 100 labels each was tedious, and homework sometimes didn't get done. To reduce friction, the team looked at third-party labeling vendors. They found tools either too use-case-specific or too general, and rather than spend weeks evaluating them, they built their own with Claude Code and Streamlit. Koh said they essentially one-shotted it.

He valued that the tool was low maintenance and low risk. If it broke, it could be fixed right away. Deploys took seconds. Non-engineers could personalize it themselves by vibe coding. This was done with Opus 4, and Koh said he expects it to be even easier with Opus 4.6. His takeaway was that building a one-off internal tool like this is sometimes easier and cheaper than buying one.

The ground-truth dataset enabled fast iteration. The team could hypothesize that employee levels were needed, add them, and run against the dataset to see whether the agent now caught the relevant cases. Koh called this a key point in development. It gave the team early confidence the product could work, which helped win internal buy-in and bring customers on as design partners.

Evals: Start With Five

On evals, Koh urged starting early and not letting perfectionism get in the way. You don't need a thousand data points. Ramp started with five cases the team knew should never fail, and kept adding. His other requirements:

  • Evals should be easy to run, a single command anyone can execute.
  • Results should be easy to understand at a glance.
  • Ideally, they should run in CI so everyone can merge code with confidence.

His reasoning: when you think you're helping an LLM or agent by adding context or tools, "more likely than not" there will be a bad consequence you didn't foresee. The context may be wrong, tool instructions may be off, or a docstring may be confusing or conflicting.

Koh also recommended online evals alongside offline ones built on historical data. He acknowledged online evals can be more confusing and harder to measure, but said any metric captured as users interact with the system is worth having as a leading indicator. For Ramp, one such metric was the rate of each decision type, especially "unsure" decisions, where the agent lacked enough information. It is a simpler eval, but it gave a useful health check on the running system.

Evals also make model upgrades safer. When new models such as Opus 4.6 or GPT-5.3 arrive, Koh said, a newer model might fix part of the problem, but it could also make things worse without any prompt or system changes. Benchmarking against evals lets the team switch models with confidence.

Letting Finance Teams Edit Their Own "CLAUDE.md"

Now that the policy agent is available across the Ramp platform, Koh shared lessons about customers. Engineers enjoy the control Claude Code gives them through editing their CLAUDE.md file. Finance people, it turns out, like the same thing, and for them the equivalent is the expense policy. When the agent makes a wrong decision, Ramp tells customers to update their policy doc.

That was initially a scary idea for finance teams, since a policy document is something "you don't mess with" and changing it normally involves many hoops. But once customers saw a fast feedback loop ("change that, you'll see it right away"), Koh said they became excited to do it.

Trust was built gradually. Ramp started with large enterprise customers, including Fortune 500 companies, reasoning that they have the most expenses and spend the most time reviewing things like coffee purchases. The agent took no autonomous action at first. Its output was framed purely as "suggestions." Eventually customers asked to move from suggestions to auto-approvals, for example: "Anything under $200, you guys are mostly right... let me just go auto approve it." Ramp responded with an "autonomy slider" customers can adjust themselves.

Koh's last point drew an analogy to LLMs. Models do well when they can test their own code and iterate, and users likewise thrive on in-product feedback loops. Giving them ways inside the product to improve the policy doc and the agent's behavior makes them eager to take ownership and personalize it.

The Applied AI Service

Ian Tracey turned to how Ramp gets leverage for its own engineers and cross-functional partners. He said the section was deliberately titled "infrastructure and culture" because both are hard problems, and changing how people work is a big part of the story.

On infrastructure, most applied AI at Ramp runs through an internal applied AI service. From a distance it resembles an LLM proxy or something like LiteLLM, but Tracey described three main extensions:

  1. Structured output and consistent APIs/SDKs across model providers. This is tricky given how fast provider APIs change, but Ramp doesn't want product teams to deal with it. Switching from GPT-5.3 to Opus, or trying Gemini 3 Pro, should be a config change.
  2. Batch processing and workflow handling. This is useful for evals and bulk document or data analysis. The service handles batching, rate limits, and whether to run jobs online or offline, so teams can focus on customer value.
  3. Cost tracing across teams and products. This lets Ramp map the Pareto curve of model performance versus cost, track how it changes over time, and spot teams building things that won't be sustainable long-term.

Tracey added that Ramp jokes its customers may be using a more frontier model than they know exists. When a new model comes out, adopting it is a one-line config change that propagates to every downstream SDK. Teams don't have to learn a new SDK or update dozens of call sites to get models Ramp has already vetted.

The Tool Catalog

Because Ramp's product handles sensitive data and workflows, Tracey said engineers often raise concerns about hallucination and safety. The speakers believe this "all comes down to the catalog of tools" teams build and integrate with. He showed Ramp's internal tool catalog, with examples like getting a policy snippet, a per diem rate, or recent transactions. These tools are built together with product teams so they capture the nuances of the data and use cases.

The catalog makes gaps visible, showing where a tool doesn't yet exist. Its tools are also usable both in internal repos and in the core product. Someone with an idea for, say, a reimbursement agent can see which tools exist, how to integrate them, and which systems they connect to, then prototype a new vibe-coded product surface without building tools from scratch. Tracey said Ramp has "many hundreds" of these tools today and, echoing Koblov, thinks it could reach multiple thousands.

Ramp Inspect: A Background Coding Agent

Tracey said Ramp noticed that its engineers had a context problem much like the one its customers have. Even with tools like Claude Code or Codex, the work engineers actually do is fragmented across many systems:

  • logs in Datadog,
  • a production database,
  • alerting systems and incident.io,
  • Slack messages,
  • Notion docs,
  • tacit knowledge held by specific product teams.

At the end of last year, Ramp set out to integrate all of this into its own background coding agent, Ramp Inspect. It can work autonomously while people are in meetings or as bugs come up. Tracey said that this month Ramp Inspect accounts for more than 50% of the PRs Ramp merges to production.

Ramp tracks usage on a dashboard, partly to create "subtle healthy competition" and partly to show people they can use it. Engineering leads in sessions by a wide margin, but product, design, risk, legal, corporate finance, marketing and CX teams also use it. Their work includes copy changes, logic fixes, and responses to incidents and bugs.

Tracey walked through a session and several design principles:

  • Sandboxed environment. Each session spins up a fast Modal code sandbox in the background. Containers can be resumed, spun up and spun down in isolation, with the same environment an engineer would have when developing at Ramp.
  • Integrated context. The agent follows a series of tasks to stay on track, creates a GitHub branch, and connects to context documents, Datadog, and a read replica so it can write queries.
  • Multiplayer first. Tracey called this a subtle design choice that turned out to have a big impact. When an engineer pairs with a designer or PM in a session, they can help that person improve their prompting, and collaborators can flag unexpected failures. That made it a real source of cross-functional collaboration.
  • Multiple entry points. Sessions can start from a Kanban UI, an API, or a Slack thread. When started from Slack, the agent takes the thread's full context, so users don't have to re-prompt with earlier conversation.
  • Full-stack capability. Ramp runs VNC inside the Modal sandbox, giving a full VS Code environment with Chrome DevTools and MCP.
  • Test access. The agent can use Ramp's 150,000+ tests, respond to CI in GitHub, and patch failures before notifying the user that the PR is ready.

Tracey pointed to builders.ramp.com, where Ramp has published the blueprint for building this, and mentioned an open-source implementation on GitHub called Open Inspect.

Team A vs. Team B

With more than half of merged PRs going through Ramp Inspect, and low-level firefighting, small fixes and tweaks spread across the company, Tracey said Ramp is rethinking how its engineering teams operate. He offered a thought experiment comparing two teams.

Team A:

  • cares about impact,
  • handles ambiguous problems,
  • understands the product, business and data,
  • adopts new tools,
  • finds creative solutions,
  • obsesses over user experience.

Team B:

  • debates libraries,
  • adds process when things feel chaotic,
  • constantly complains about headcount,
  • bikesheds details (functional programming or not, which TypeScript library) instead of focusing on the user,
  • builds before understanding the problem ("we're going to just vibe code this, bro"),
  • fixates on performative code quality and subjective nitpicks.

Tracey said he has worked on both kinds of teams and argued they will diverge.

He cited a Harvard study from around the end of last year on hiring trends for junior and senior engineers since AI tools accelerated. In his view, the study treats this as a years-of-experience issue, but the real dividing line is the set of Team A qualities. Coding "was never really the hardest part" of many jobs. Staff and staff-plus engineers are paid for judgment, context, the ability to see around corners, and "scar tissue." If you ask Opus 4.6 to do something, those engineers know when the result won't work or is a bad idea.

Tracey said media narratives about coding agents miss that "you could still build the wrong thing just a lot faster," and create bigger messes. The skills he expects to matter more are:

  • figuring out what to build and understanding users well enough,
  • selling ideas to skeptical stakeholders,
  • making good design decisions with incomplete information,
  • maintaining momentum through a project's "long middle."

On selling ideas, he noted it was not obvious that Ramp should spend time building a background coding agent. On the long middle, he connected it to current debates about SaaS and the stock market. Vibe coding something is easy, he said, but getting through that middle to a deployed product with product-market fit is why good engineers are still needed, and "not enough people recognize that."

Where This Leaves Engineering

Tracey acknowledged the "doomerism and scariness" around AI narratives but said he sees this as an exciting time to build. Unlike factory work or farming, he said, software is never done, and he referenced Ramp's internal "job's not finished" meme. He rejected the idea that if everyone is twice as productive, companies need half the people. With freed-up capacity, he predicted four things:

  1. Companies will chase opportunities they couldn't previously afford. He said he's not sure Ramp would be pursuing agentic workflows and bigger problems in the financial stack without this technology.
  2. Companies will enter adjacent markets and stitch together more value for customers.
  3. Teams will rebuild systems that were too expensive to touch. Building an internal background coding agent at a financial operations software company once seemed like a crazy idea and now makes sense.
  4. The bar for "good enough" will rise. He expects building more "mind-blowing" experiences and delivering more value to be the story of the next decade.