How Ramp Moved From Many AI Agents to One Agent With Many Skills
The Pragmatic EngineerAt The Pragmatic Summit, four Ramp engineering leaders described what the company learned building AI products for finance teams: Nik Koblov (EVP Engineering), Veeral Patel (Director, Applied AI and Spend), Will Koh (Staff Engineer, Applied AI), and Ian Tracey (Staff Software Engineer). The session covered four things:
- A strategic pivot away from building many separate agents.
- A detailed account of how Ramp built its policy agent.
- The internal infrastructure behind Ramp's applied AI work.
- An argument about how engineering culture has to change when coding agents take over more routine work.
Koblov framed the talk around what the speakers see as a paradigm shift in software that "requires complete rethink," and, with it, a simplification of the stack.
The Coffee Problem: Where Ramp Started With AI
Koblov introduced Ramp as a finance platform for modern businesses with more than 50,000 customers, whose business is "saving you time and money." To illustrate, he used a cup of coffee bought on a company card. By his account, the purchase usually costs an employee about 15 minutes, because of three simple steps that each take minutes. That cost compounds across a company.
Ramp handles that whole chain agentically, Koblov said:
- writing a memo after the card is tapped,
- classifying the transaction against the general ledger,
- sourcing and attaching the receipt,
- normalizing the merchant against the company's merchant inventory.
This was Ramp's first use of AI, starting roughly three years ago with "one-shot" tasks such as normalizing a merchant or writing a memo. Koblov said these have worked "really, really well as the models get better."
From there, the ambition grew. Koblov argued that nearly every finance-related role is losing time to manual work, including AP clerks, finance teams, purchasing teams and data teams. The complexity "has a ramp shape," increasing as you move through different jobs to be done. His example was a Slack channel called "help data," where someone would ask for a CSV and "a poor person will go and write a SQL query." Ramp replaced that process about a year and a half ago. The company's end-state goal is to cover everything admins, employees and finance teams do that is not directly related to making money.
The Pivot: One Agent, a Thousand Skills
The first major lesson Koblov described was organizational as much as technical. Last year, Ramp intentionally let each team experiment on its own. The result was "maybe four different ways of doing the same thing," for both synchronous agents and background agents. His conclusion: "you don't need to build a thousand agents." Instead, a company should drive its framework toward "a single agent with a thousand skills."
To explain why, Koblov broke any modern AI-driven process into five parts:
- an event, such as receiving an invoice that needs paying;
- prompt instructions describing what to do, plus guardrails such as an expense or payables policy;
- the context the agent should consider;
- tools;
- the APIs and actions that can be taken.
In his framing, traditional software focused only on the last two, the tools and actions. In the new paradigm, "software is doing everything." The goal becomes an autonomous system that can react, reason and act with little or no human supervision.
In practice, the first step was consolidating conversational interfaces. At the end of last year, Koblov said, Ramp had about five different conversational UXs. These have been merged into one, called OmniChat ("Omni" meaning omnipresent), which is being deployed to every surface of the product. He stressed that it coexists with traditional UI, because "you still need tables and buttons" and people don't always want to talk to their software.
His demo was a request to onboard a new employee. OmniChat resolved the person to an employee ID and looked up their place in the corporate structure through an HRIS tool. It then found a previously created agentic workflow called the "new hire playbook" and asked whether to use it.
This runs on an in-house, lightweight agent framework that orchestrates tools. Koblov said engineers build these tools very quickly. More recently, one product manager vibe-coded about 20 tools, so engineers are "no longer needed to build these tools."
Playbooks handle more involved workflows. For onboarding, a user can describe in plain language what should happen when someone joins:
- give them a card,
- make sure they submit receipts for every transaction,
- congratulate them on Slack,
- check in after two weeks.
Ramp compiles that description into a runnable, deterministic workflow and hands it to the agent to execute.
Koblov then showed how the pieces connect when a card is swiped:
- A real-time policy review runs, enforcing the company's spend requirements. He said this is what makes it safe to give cards to every employee.
- A handoff goes to an accounting-code agent that classifies the transaction using the finance team's rules. Traditional products would expose that choice to employees, who usually don't know how a transaction maps to the GL. The agent does it better, Koblov argued, because it has the full chart of accounts and understands the ERP.
- The agent either auto-approves or, in the worst case, brings in a human to review for materiality or to flag out-of-policy spend.
What the Policy Agent Does
Veeral Patel took over to go deeper on the policy agent. Finance teams look at receipts every day, sometimes hundreds or thousands of them, and Patel admitted he would likely make mistakes judging them by eye. He walked through three examples of the agent's decisions:
- A dinner receipt. The agent reasoned over the image and the transaction data and found eight guests on the receipt, which Patel said he could barely see himself. It confirmed the cost was under Ramp's internal $80-per-person cap and that the occasion was a team welcome dinner. With the amount and merchant verified, it recommended approval.
- An OpenAI charge. A colleague had been testing ChatGPT features. The agent judged this a valid business expense and recommended approval.
- A $3 bakery charge. The agent rejected it because it wasn't part of an overtime purchase and didn't happen on a weekend.
The product grew out of a customer request. A Fortune 500 customer asked Ramp to approve certain types of expenses and reject others, and brought a list of rules. Patel had worked on early versions of Ramp's deterministic rules. He said the team chose not to keep adding incremental rules. Instead, they took a page from Andrej Karpathy's line that English is the new programming language and turned the expense policy document itself into the rules. He showed Ramp's own expense policy next to a production screenshot and said the product is seeing "really great" use.
Building It Like an Early-Stage Startup
Patel described running the project like an early-stage startup inside Ramp. The team found design partners, including that Fortune 500 company, and held weekly meetings with each of them to collect feedback and decide what to improve.
The main cultural lesson, in Patel's view, was accepting that "AI products cannot be one-shotted." Everyone on the team (PMs, designers, engineers) had to agree that the product would not be perfect on day one.
The team dogfooded internally and began with an even narrower problem: whether a "coffee with a colleague" expense should be approved. Ramp's finance team considered these low-risk, single-dollar-amount transactions.
An early finding once the agent reached production was that wrong decisions stemmed less from the models and more from the context the models were given. Patel said the team could have tried to list all the needed context before starting engineering. They decided it was better to learn from live internal data. One example: an employee's role and title matter a lot when reading a policy, because C-suite employees may have higher limits or be allowed first-class flights on certain routes. So the team began extracting more information from receipts and pulling in HRIS fields already available in Ramp.
From a Simple Pipeline to an Agentic Black Box
Will Koh described how the architecture evolved. The team initially "went big," aiming to automate all of finance and all reviews. In practice, they started small with questions like whether a cup of coffee fits the expense policy. Koh explained that even a simple-sounding question grows complex. Even if the team had gotten the context design right up front, he expected it would still break once generalized to other businesses. His principle: the simpler the system, the easier it is to iterate on, and complexity can be layered on once you know what works.
The first version was a classic pipeline. An expense comes in, the system retrieves context, and a series of well-defined LLM calls asks whether the expense is in policy, why, and how to show that to the user. The output is then formatted for the user.
The next stage recognized that expenses differ. The system classified each one (travel, meal, entertainment), used conditional prompting, retrieved context based on the category, and gave the model some tools. That let it decide on its own that it needed, for example, flight information or the employee's level.
A few iterations later, the result was a full agentic workflow. The agent has complex read tools across the platform, drawn from a company-wide internal toolbox that all of Ramp's agents share. It also got write capabilities: it now records decisions, writes its reasoning, and auto-approves expenses on users' behalf, all in a loop.
Koh was explicit about the trade-off. As a system goes from simple to complex, capability and autonomy go up and the AI "seems smarter," but traceability and explainability go down. The team can read the model's reasoning tokens, but "in the end, we have no control over it." A small black box becomes a bigger one.
His response was to build strong auditability from the start. Assume you know only the inputs and outputs, and ask whether you can verify the agent did the right thing from those alone, even as the black box changes.
Users Aren't Ground Truth
Koh said Ramp initially assumed, as with many products, that users' own decisions would be correct: if a user approved, the agent should approve. It turned out users are often wrong. They may not know the policy, they may trust their employees, they may be lazy, or "it's a Sunday." When the agent copied user behavior, finance teams would come back and say an expense shouldn't have been on the company card.
So Ramp had to define its own standard of correctness. The team held a weekly cross-functional labeling session, which Koh said produced two benefits:
- A ground-truth dataset the team trusted and could always test against.
- Shared understanding. When the agent got something wrong or was missing context, everyone knew. That reduced communication overhead and aligned priorities.
The sessions were expensive, though. Getting people in a room weekly and assigning homework of 100 labels each was tedious, and homework sometimes didn't get done. To reduce friction, the team looked at third-party labeling vendors. They found tools either too use-case-specific or too general, and rather than spend weeks evaluating them, they built their own with Claude Code and Streamlit. Koh said they essentially one-shotted it.
He valued that the tool was low maintenance and low risk. If it broke, it could be fixed right away. Deploys took seconds. Non-engineers could personalize it themselves by vibe coding. This was done with Opus 4, and Koh said he expects it to be even easier with Opus 4.6. His takeaway was that building a one-off internal tool like this is sometimes easier and cheaper than buying one.
The ground-truth dataset enabled fast iteration. The team could hypothesize that employee levels were needed, add them, and run against the dataset to see whether the agent now caught the relevant cases. Koh called this a key point in development. It gave the team early confidence the product could work, which helped win internal buy-in and bring customers on as design partners.
Evals: Start With Five
On evals, Koh urged starting early and not letting perfectionism get in the way. You don't need a thousand data points. Ramp started with five cases the team knew should never fail, and kept adding. His other requirements:
- Evals should be easy to run, a single command anyone can execute.
- Results should be easy to understand at a glance.
- Ideally, they should run in CI so everyone can merge code with confidence.
His reasoning: when you think you're helping an LLM or agent by adding context or tools, "more likely than not" there will be a bad consequence you didn't foresee. The context may be wrong, tool instructions may be off, or a docstring may be confusing or conflicting.
Koh also recommended online evals alongside offline ones built on historical data. He acknowledged online evals can be more confusing and harder to measure, but said any metric captured as users interact with the system is worth having as a leading indicator. For Ramp, one such metric was the rate of each decision type, especially "unsure" decisions, where the agent lacked enough information. It is a simpler eval, but it gave a useful health check on the running system.
Evals also make model upgrades safer. When new models such as Opus 4.6 or GPT-5.3 arrive, Koh said, a newer model might fix part of the problem, but it could also make things worse without any prompt or system changes. Benchmarking against evals lets the team switch models with confidence.
Letting Finance Teams Edit Their Own "CLAUDE.md"
Now that the policy agent is available across the Ramp platform, Koh shared lessons about customers. Engineers enjoy the control Claude Code gives them through editing their CLAUDE.md file. Finance people, it turns out, like the same thing, and for them the equivalent is the expense policy. When the agent makes a wrong decision, Ramp tells customers to update their policy doc.
That was initially a scary idea for finance teams, since a policy document is something "you don't mess with" and changing it normally involves many hoops. But once customers saw a fast feedback loop ("change that, you'll see it right away"), Koh said they became excited to do it.
Trust was built gradually. Ramp started with large enterprise customers, including Fortune 500 companies, reasoning that they have the most expenses and spend the most time reviewing things like coffee purchases. The agent took no autonomous action at first. Its output was framed purely as "suggestions." Eventually customers asked to move from suggestions to auto-approvals, for example: "Anything under $200, you guys are mostly right... let me just go auto approve it." Ramp responded with an "autonomy slider" customers can adjust themselves.
Koh's last point drew an analogy to LLMs. Models do well when they can test their own code and iterate, and users likewise thrive on in-product feedback loops. Giving them ways inside the product to improve the policy doc and the agent's behavior makes them eager to take ownership and personalize it.
The Applied AI Service
Ian Tracey turned to how Ramp gets leverage for its own engineers and cross-functional partners. He said the section was deliberately titled "infrastructure and culture" because both are hard problems, and changing how people work is a big part of the story.
On infrastructure, most applied AI at Ramp runs through an internal applied AI service. From a distance it resembles an LLM proxy or something like LiteLLM, but Tracey described three main extensions:
- Structured output and consistent APIs/SDKs across model providers. This is tricky given how fast provider APIs change, but Ramp doesn't want product teams to deal with it. Switching from GPT-5.3 to Opus, or trying Gemini 3 Pro, should be a config change.
- Batch processing and workflow handling. This is useful for evals and bulk document or data analysis. The service handles batching, rate limits, and whether to run jobs online or offline, so teams can focus on customer value.
- Cost tracing across teams and products. This lets Ramp map the Pareto curve of model performance versus cost, track how it changes over time, and spot teams building things that won't be sustainable long-term.
Tracey added that Ramp jokes its customers may be using a more frontier model than they know exists. When a new model comes out, adopting it is a one-line config change that propagates to every downstream SDK. Teams don't have to learn a new SDK or update dozens of call sites to get models Ramp has already vetted.
The Tool Catalog
Because Ramp's product handles sensitive data and workflows, Tracey said engineers often raise concerns about hallucination and safety. The speakers believe this "all comes down to the catalog of tools" teams build and integrate with. He showed Ramp's internal tool catalog, with examples like getting a policy snippet, a per diem rate, or recent transactions. These tools are built together with product teams so they capture the nuances of the data and use cases.
The catalog makes gaps visible, showing where a tool doesn't yet exist. Its tools are also usable both in internal repos and in the core product. Someone with an idea for, say, a reimbursement agent can see which tools exist, how to integrate them, and which systems they connect to, then prototype a new vibe-coded product surface without building tools from scratch. Tracey said Ramp has "many hundreds" of these tools today and, echoing Koblov, thinks it could reach multiple thousands.
Ramp Inspect: A Background Coding Agent
Tracey said Ramp noticed that its engineers had a context problem much like the one its customers have. Even with tools like Claude Code or Codex, the work engineers actually do is fragmented across many systems:
- logs in Datadog,
- a production database,
- alerting systems and incident.io,
- Slack messages,
- Notion docs,
- tacit knowledge held by specific product teams.
At the end of last year, Ramp set out to integrate all of this into its own background coding agent, Ramp Inspect. It can work autonomously while people are in meetings or as bugs come up. Tracey said that this month Ramp Inspect accounts for more than 50% of the PRs Ramp merges to production.
Ramp tracks usage on a dashboard, partly to create "subtle healthy competition" and partly to show people they can use it. Engineering leads in sessions by a wide margin, but product, design, risk, legal, corporate finance, marketing and CX teams also use it. Their work includes copy changes, logic fixes, and responses to incidents and bugs.
Tracey walked through a session and several design principles:
- Sandboxed environment. Each session spins up a fast Modal code sandbox in the background. Containers can be resumed, spun up and spun down in isolation, with the same environment an engineer would have when developing at Ramp.
- Integrated context. The agent follows a series of tasks to stay on track, creates a GitHub branch, and connects to context documents, Datadog, and a read replica so it can write queries.
- Multiplayer first. Tracey called this a subtle design choice that turned out to have a big impact. When an engineer pairs with a designer or PM in a session, they can help that person improve their prompting, and collaborators can flag unexpected failures. That made it a real source of cross-functional collaboration.
- Multiple entry points. Sessions can start from a Kanban UI, an API, or a Slack thread. When started from Slack, the agent takes the thread's full context, so users don't have to re-prompt with earlier conversation.
- Full-stack capability. Ramp runs VNC inside the Modal sandbox, giving a full VS Code environment with Chrome DevTools and MCP.
- Test access. The agent can use Ramp's 150,000+ tests, respond to CI in GitHub, and patch failures before notifying the user that the PR is ready.
Tracey pointed to builders.ramp.com, where Ramp has published the blueprint for building this, and mentioned an open-source implementation on GitHub called Open Inspect.
Team A vs. Team B
With more than half of merged PRs going through Ramp Inspect, and low-level firefighting, small fixes and tweaks spread across the company, Tracey said Ramp is rethinking how its engineering teams operate. He offered a thought experiment comparing two teams.
Team A:
- cares about impact,
- handles ambiguous problems,
- understands the product, business and data,
- adopts new tools,
- finds creative solutions,
- obsesses over user experience.
Team B:
- debates libraries,
- adds process when things feel chaotic,
- constantly complains about headcount,
- bikesheds details (functional programming or not, which TypeScript library) instead of focusing on the user,
- builds before understanding the problem ("we're going to just vibe code this, bro"),
- fixates on performative code quality and subjective nitpicks.
Tracey said he has worked on both kinds of teams and argued they will diverge.
He cited a Harvard study from around the end of last year on hiring trends for junior and senior engineers since AI tools accelerated. In his view, the study treats this as a years-of-experience issue, but the real dividing line is the set of Team A qualities. Coding "was never really the hardest part" of many jobs. Staff and staff-plus engineers are paid for judgment, context, the ability to see around corners, and "scar tissue." If you ask Opus 4.6 to do something, those engineers know when the result won't work or is a bad idea.
Tracey said media narratives about coding agents miss that "you could still build the wrong thing just a lot faster," and create bigger messes. The skills he expects to matter more are:
- figuring out what to build and understanding users well enough,
- selling ideas to skeptical stakeholders,
- making good design decisions with incomplete information,
- maintaining momentum through a project's "long middle."
On selling ideas, he noted it was not obvious that Ramp should spend time building a background coding agent. On the long middle, he connected it to current debates about SaaS and the stock market. Vibe coding something is easy, he said, but getting through that middle to a deployed product with product-market fit is why good engineers are still needed, and "not enough people recognize that."
Where This Leaves Engineering
Tracey acknowledged the "doomerism and scariness" around AI narratives but said he sees this as an exciting time to build. Unlike factory work or farming, he said, software is never done, and he referenced Ramp's internal "job's not finished" meme. He rejected the idea that if everyone is twice as productive, companies need half the people. With freed-up capacity, he predicted four things:
- Companies will chase opportunities they couldn't previously afford. He said he's not sure Ramp would be pursuing agentic workflows and bigger problems in the financial stack without this technology.
- Companies will enter adjacent markets and stitch together more value for customers.
- Teams will rebuild systems that were too expensive to touch. Building an internal background coding agent at a financial operations software company once seemed like a crazy idea and now makes sense.
- The bar for "good enough" will rise. He expects building more "mind-blowing" experiences and delivering more value to be the story of the next decade.
Today we're going to talk about AI at Ramp, and I'm going to give an intro, a quick introduction into what Ramp is. Really briefly, we're going to walk through the simplest possible expense use case that you guys can all resonate with, cuz I see everybody's drinking coffee.
And then we're going to talk quickly about a lesson that we learned this year while we were building Brazilian agents, and sort of the pivot in the paradigm that's happening, especially after February 6th. And then we're going to double click onto how we built one of our most popular agents, the policy agent. And then finally, we'll dig in into the infrastructure built that this is requiring us to do on our side. And in my mind, most importantly, the culture shift that needs to happen on everyone's teams in order to be able to operate in a way that delivers products into the hands of your customers in the fastest and most impactful way.
So without further ado, quick intro about Ramp. We are the number one finance platform for modern businesses with 50,000 plus customers, and we're in the business of saving you time and money. I've seen some of those names on the name tags here. So thank you for being Ramp customers.
Really exciting. Really quickly, so a cup of coffee usually takes about 15 minutes of your time, cuz you got to do these three simple things which unfortunately take minutes. This compounds through the company. And what Ramp does, in the simplest possible way, we just condense time and return money back.
So a simple story over a transaction, from tapping the card to writing a memo to classifying the transaction according to your GL to sourcing the receipt, attaching the receipt, normalizing the merchant to your inventory of merchants, is all done agentically at Ramp. And this was our first foray, probably by now — you guys still feel here? Yeah. Probably by now about 3 years ago, we started doing these one-shot things with AI: normalize merchant, write a memo. And it's been working really, really well as the models get better.
What else is going on at the company? Well, literally every persona at the company is wasting time on a lot of manual work. So from AP clerks to your finance team, from your purchasing teams, keep going to more finance work, your data teams. At Ramp we used to have a channel called help data where somebody will ask for a CSV and a poor person will go and write a SQL query. We replaced it about a year and a half ago.
So a lot of time is being spent, and the complexity has a ramp shape. It only increases as you go through different jobs to be done. So if you guys watched the Super Bowl, you might be familiar with Brian, our agent. So we've been writing a lot of agents, literally for every job to be done, to cover, in the end state, the entirety of what admins, employees, and finance teams are doing that is not directly related to making the money. We want you all to be making money and focus on your customers, not on how to close the books.
But what's been happening for the past few weeks is that we're living through the most exciting paradigm shift in software, and it requires a complete rethink, and with rethink, simplification of your stack. So, what we learned is you don't need to build a thousand agents. We intentionally last year allowed each individual team to go and experiment, and we ended up with maybe four different ways of doing the same thing, both for synchronous agents as well as for background agents. But instead you want to drive your framework towards a single agent with a thousand skills.
So, let's talk about what software traditionally used to focus on. Every process, especially in the modern AI stack, boils down to having an event. So, a prompt: you can receive an invoice and you want to pay it. Some prompt instructions of what you want to do with it, and some guardrails, like a policy, like an expense policy or your payables policy. Context: what is the data that the agent should consider. And then finally, tools. These are APIs and actions that you can do. And traditionally software would focus on only four and five.
In the new paradigm, software is doing everything. So, you want to focus on building an autonomous system of action that can react, reason, and act without a human or with very little human supervision.
So, what does it mean in terms of what we're building? So, first, we decided we were going to consolidate the interactions, verbal interactions with the agents, to a single conversational UX. We literally, at the end of last year, had about five different conversational UXs. We have now consolidated it into what we call an OmniChat, Omni meaning omnipresent. It is now being deployed to every surface of the product. And it works well with the traditional UX, because you still need tables and buttons, and you don't always want to be talking to your software.
But this is a good example of what OmniChat looks like. "Please onboard a new employee." OmniChat can resolve an employee to an employee ID and look up, through an HRIS tool, their corporate structure, and it found an agentic workflow that we created previously called the new hire playbook. And the agent is asking, would you like me to onboard the person using this playbook?
How is this possible? We built an in-house lightweight agent framework that provides orchestration with tools that engineers are very quickly building, and most recently we had one product manager vibe-code about 20 tools, so engineers are no longer needed to build these tools.
And sometimes your workflows are involved, such as employee onboarding, which consists of four steps. So you can just go into Ramp and describe what you want to happen when a new employee joins: give them a card, make sure they get receipts for every transaction, congratulate them on Slack, and check in with them in two weeks. We are now able to compile this into a runnable deterministic workflow and then give it to the agent to execute. Playbooks make use of tools.
And how this all comes together — this is an example which Veeral is going to double click on next — is upon swiping the card, there's a real-time policy review that's happening directly in the software, and the policy agent enforces your company requirements with regard to spend. Therefore, it's very safe to give Ramp cards to literally every employee in your company. And there's a handoff happening with an accounting coding agent that classifies this transaction and applies the rules of your back office team, of your finance team. As an employee, I have no idea how certain transactions should match to our GL, and that's what typical traditional products would do: they would expose it to you. So the agent is much better at doing it, because it has the full context of your chart of accounts, it understands your ERP, and then it can either auto-approve or, in the worst-case scenario, it will involve the human in the loop to review by materiality, or notify that there is an out-of-policy spend.
With that, please welcome Veeral, who will dive deeper into the policy agent.
Thanks, Nik.
Oops. Awesome. So, a lot of finance teams are looking at receipts like this basically every day, and maybe they might have hundreds or thousands of these. If you told me to look at this and decide if I should approve or reject this transaction, I'm probably going to make a mistake.
So, policy agent basically reasons on this image and all the transaction data that we have, and told me that there were eight guests on the receipt. I could barely see that when I was looking at it. It was below the $80 a person cap that we have internally. They were going for a team welcome dinner. And so because the amount was verified as well, and the merchant, policy agent told me to approve this transaction. Similarly, for this OpenAI transaction, Anan was testing out some ChatGPT features, and so policy agent told me this was a valid business expense and told me to approve it. And then this $3 bakery charge was rejected because it wasn't part of an overtime purchase and it didn't happen on the weekend.
So, really we looked at this as an opportunity to rethink how Ramp was set up. Controllers and finance teams are looking at transactions like these and making these decisions every day. And a Fortune 500 company that is one of our customers was coming to us and saying, "Hey, can you make sure that you approve these types of expenses and reject these types of expenses?" And they basically had a list of all the rules that Ramp should follow. And we kind of saw this as an opportunity not to add more incremental deterministic rules that kind of define our product — and I worked on some of the first versions of these — but actually to take a page from Andrej Karpathy, saying that English is the new programming language, and kind of turn the expense policy into the rules themselves.
You can see Ramp's expense policy on the left, and this is a screenshot from our production environment, but we are seeing really great use out of our policy agent product.
And it kind of needed to start really organically. So, we kind of operated like an early-stage startup. We're already very incremental and fast at Ramp, but we found some design partners, like that Fortune 500 company. We iterated really quickly, and we had weekly meetings with all of them to understand exactly what feedback we wanted to hear and what we could improve.
I think one of the main important things that we realized across Ramp is that we really needed to lean into the fact that AI products cannot be one-shotted. You need to start with something simple. And so, as long as everyone on your team, PMs, designers, engineers, is aligned that you're not going to have perfection on day one — I think that was actually one of the main cultural learnings.
And so, we dogfooded a lot of this work internally and started with an even more constrained problem of trying to decide whether our coffee with a colleague transactions should be approved or rejected. These are single dollar amount transactions that are low risk according to our finance team. And so, we started with these transactions.
And one of the early learnings, especially as we released this into production, was that a lot of the reason that policy agent would be wrong would be less about the models themselves and more about the context that we were giving to the LLMs themselves. So, we could have sat down and thought about all the context in the beginning, before we even kicked off any engineering work, but we realized actually the best thing would be to learn from some of our live internal data. And so, for example, we learned that the role and the title of an employee are super important when looking at expense policy docs. The C-suite, for example, might have higher limits; maybe they can fly first class for certain flights. And so, we started extracting more information from receipts and started pulling in information from HRIS fields that are already on Ramp.
And so, Will is going to talk you through exactly the iterations that we went through to implement policy agent and some of the learnings along the way.
Is this down? It's down. Yeah. Okay.
All right, cool. Awesome. So, when we first started building the policy agent internally, we dreamed, we went big. We're like, "Hey, let's automate all of finance. Let's automate all reviews." But when it came down to it, we actually had to start small: is that cup of coffee, you know, in your expense policy?
And the reason that we did that was because even though the problem sounds simple to automate — you know, it's a simple question, is this in policy or not? — it was going to grow to be complex. Kind of like Veeral said, we could have gone down and figured out what context do we have, how can we add it, how can we put it all together in a way that an LLM can understand, and put it all together from the get-go. But we knew that even if we aimed and got everything right the first time, it was probably going to be wrong once you applied and generalized it to another business.
So, the simpler the system, I think the easier it is to iterate on top of it. And once you iterate, you know what's going to work, you know what's not, and you can layer complexity on top of that. And I think that's pretty important to keep in mind when you're building an LLM or an agent starter.
So, for us, we started really simple, kind of the classic: we have an expense come in, retrieve the context around it, pass it through a series of LLM calls that are very well defined, like, "Hey, is this in policy? Why is it in policy? How can we show the user that it's in policy?" And then give an output that makes sense in this way to the user.
Eventually, we learned that each expense is kind of different. We can classify an expense based on: is it travel? Is it a meal? Is it entertainment? Do conditional prompting, and then retrieve context based on that, pass it through a series of LLM calls, and give it some tools so that it can also autonomously decide, "Hey, I need flight information actually," or "I need this employee's level," and kind of layer that on top.
And a few iterations later, we came to a full-on agentic workflow. We ended up with complex tools to read across all of our platform, and these tools are shared across all of our agents. It's not just for policy agent. We have a company internal toolbox that all of our agents can easily reach into and use. And we gave it the capability to write as well. So, it's now writing decisions, it's writing reasoning, it's auto-approving expenses on users' behalf. And it goes in a loop. So, now it's more of a black box, and that's kind of the trade-off you get.
As you go from simple to complex systems, your capability goes up, your autonomy goes up, your agents are able to do more, your AI can do more, your AI seems smarter. But in exchange, you're losing traceability and explainability. We look at it now, we can kind of look at the reasoning tokens that the LLM gives us, but in the end, we have no control over it. It's going to do what it thinks is right, it's going to make the tool calls, it's going to tell you it's right or wrong. So, a smaller black box becomes a bigger black box as the system becomes more complex.
So, one thing that is really important when doing something like this is that from the beginning, you need really good auditability. Even if you know how it works, assume that your inputs and outputs are all you know, and make sure that it's correct. So if it was a black box system and you only saw the input and output, can you verify that it did the right thing? And even if that black box changes, you should be able to reason about whether the output is correct.
As with many products that we built at Ramp and across other companies, we thought that the users would be correct. You know, if the user says approve, the agent should approve. If the user says reject, the agent should reject. But it turns out the users are actually incorrect. They're wrong sometimes. They don't know the expense policy, they trust their employees, they're lazy, it's a Sunday, who knows. So, it turns out we can't always do what the users are doing, cuz sometimes that's where the finance teams come back to you and are like, "Hey, this is wrong. This shouldn't be on the company card."
So, we have to define our own definition of correctness. And to do that, we had a weekly labeling session across the functions that are working on this product. And that had two really good outcomes. One was that we had a ground truth data set that we could always test against, and we knew that this was correct. And two was that everyone was on the same page. If our agent got
something wrong, everyone knew that it got it wrong. Or, you know, our agent is missing context, everyone knew that it's missing that context. So, there was less communication, everyone's on the same page, and they could focus on what's really priority and kind of have alignment on that.
Initially, getting all those people together in a room every week, giving them homework to label 100 data points, it's expensive. You know, everyone has things to do and sometimes they don't come back with their homework done. It just kind of almost becomes tedious even though it's so important. So, we wanted to make it as simple as possible, and the way we did that was that we looked for third-party vendors that could provide us the tools to label data and collect the data.
But, turns out some tools are too specific to a use case, some tools are too general, and we could have spent weeks trying out different tools, but we decided let's just build our own. So, we used Claude Code using Streamlit. We basically one-shotted all of this, and the greatest part of it all is that it's low maintenance, low risk. It's in our codebase; if it breaks, we can fix it right away. Deploys happen in like seconds. And non-engineers can go and personalize it. They can vibe code it. They can Claude Code it. And this was at Opus 4. So, now at Opus 4.6, I expect it's even better, and with something like that, it's definitely easier and cheaper sometimes to do something one-off like this.
And with the ground truth data set, we were able to make quick iterations. We're able to find out, "Hey, we need employee levels. Add that. How does that work?" Running it against this data set, does it actually catch it? And now say accept or approve. And we're able to make really quick iterations, and that was actually kind of a key point in developing this. We had really early confidence that this could actually work, and we were able to actually get a lot of buy-in, get a lot of customers on board and kind of try it out as a design partner.
And as part of doing that iteration with the data set, you had evals, and I feel like obviously everyone in the room now knows about evals and what they mean, but it's pretty important to have them early on. Don't let perfectionism get in the way. You don't need a full data set of a thousand data points that you're testing against every iteration. We started with five. And we knew that those five we were not going to fail. We kept adding and adding and adding.
And make sure it's easy to run. Anyone could go and just run that command. And then make sure that the results are really easy to understand. They're able to look at it, get instant output, and understand, "Hey, this is what the model's doing. This is good, this is bad," and if you want to do it as part of your CI, then everyone now can just hopefully merge in code, because whenever you think you're doing something right for the LLMs or agent, giving more context, giving it tools, more likely than not it's probably going to have some kind of bad consequence that you didn't see happening. Context was wrong, whether it be the tool instructions were wrong, or maybe the docstring was a little confusing and conflicting. So it might have consequences. You just want to make sure you're catching against those.
And then I'll touch on it briefly, but online evals are also great. So these are offline. You have a data set, it's historical, you're testing it. Online evals can be a little more confusing and harder to measure, but if you can measure anything as your users are interacting with the system, definitely as a leading metric also set them up. And for us, part of that was, hey, what are our rates of decisions. We had an unsure decision, which just meant that the agent didn't have enough information. So we can measure that online. So it's a much simpler eval, but that also gave us a pretty good health check as our system was running.
Cool. And another great part about evals is that with evals, you can make confident model changes. Whenever a new model comes out, Opus 4.6, GPT-5.3, you want to make sure that you can leverage those new models, because sometimes that could be the difference between your system getting one part of the problem wrong to right, but it could also be the opposite. It could actually be not good without any prompt changes or changing how your system works. So having evals set up and being able to benchmark really helps make confident model changes.
Cool. So now that policy agent, we've been developing this for a while, it's available for everyone on the Ramp platform. Some of the things that we learned along the way is that Claude Code as engineers is very exciting. We have full control, we get to modify our CLAUDE.md, we get to tell it to not leave comments, and it won't leave comments, hopefully.
Turns out it's not just us; finance people also really like to modify their CLAUDE.md, which is their expense policy. So, if something went wrong with the decision, then we just tell them, "Hey, go update your policy doc." To them it's a little scary concept to begin with. This is a document, you don't mess with that. You have to go through a lot of hoops if you're going to mess with that. But then it turns out if you get them really excited about the feedback loop, "Hey, change that. You'll see it right away," they'll be really excited to do this.
And then trust builds over time. So, some of the earlier customers that we had were some of the Fortune 500. We actually started with really big enterprise customers because we thought that they would have the most value. They have the most expenses coming in. They have the most time to spend on reviewing coffee expenses. So, roll it out to them. Let them build the trust. We didn't do any autonomous action. We're just like, "Hey, we're going to give you a suggestion." That's kind of how we phrased it. Suggestions.
And then eventually they came to us and were like, "Okay, you know what? I want to go from suggestions to auto approvals. Anything under $200, you guys are mostly right. I don't care about this. Let me just go auto approve it." So, we gave them the autonomy slider. We gave them a way to just turn it on, and then they actually could do it themselves.
And then, last but not least, similar to LLMs, users thrive in in-product feedback loops. So, when you're building an AI product and you have a full way for LLMs to test if their code was right and still go iterate, users are the same way. Give them in-product ways to improve the expense policy doc, improve the agent and how it operates. And they're more than excited to kind of take it over themselves and improve it and personalize it for them.
So, from here I'll pass it on to Ian, who's going to talk about the infrastructure and the culture that we have at Ramp that led us to building the policy agent.
Hey, everybody. So, you've heard a little bit about how we're getting leverage for all of the different finance teams as we operate on top of their financial infrastructure and really try to get leverage for our customers. But I think a big thing that we also spend a lot of time thinking about is how can we get leverage for Ramp itself, the engineers, our XFN works, all the people that we work with every single day. And this section is pretty intentionally named AI infrastructure and culture, because we think that this is both a really challenging infrastructure problem, but it's also a really challenging culture problem, and changing how you work as well is a big part of the story.
And so to start on the infrastructure side, the core of how most of applied AI happens at Ramp is our applied AI service. And at a 10,000 ft view, this looks something kind of like an LLM proxy or something like LiteLLM, but there's really three main extensions that we've invested in to make this a lot more powerful for a lot of our use cases.
The first is structured output and consistent API and SDKs across different model providers. This can be pretty tricky to do, especially with how quickly the APIs are changing, but it's a problem that we don't want downstream product teams to have to think about. So if you have an idea of, I want to switch from GPT-5.3 to Opus, or I want to try Gemini 3 Pro, you should be able to do that with a config change and really quickly be able to iterate on semantic similarity and trying to do a bunch of different code sandboxing and structured output calls that way.
The other thing that we've spent a ton of time thinking about is batch processing and workflow handling. This is really useful for evals or if you're doing, for us, bulk document or data analysis, and that's something that we also don't want teams to have to spend a bunch of time on: how do you want to batch this and handle it with rate limits, and do we want to do this on an offline or online job with something like Anthropic? We just want to handle that for downstream consumers so they can just focus on providing value for downstream customers.
And then the last, which is a pretty big deal, is the ability to trace different costs across teams and against products as well. And this allows us to identify the Pareto curve of what is the best model performance for cost, how are these evolving over time, what teams are actually not building something that's going to be sustainable long-term for different product services. And this can be really, really important to just remove all this work from internal teams having to think about this.
And the last thing that's kind of funny to think about, and we often joke that our customers are actually using more of a frontier model than they may even know is out yet, is it allows us to stay at the frontier. When a new model comes out, it's a one-line config change that impacts every single SDK downstream. And so, rather than teams having to learn the SDK or go into dozens of different call sites, they can just change it in one place for their specific team, and they now get the benefit of being on the latest and greatest models that we've vetted and built into the rest of the system.
Our product, as you've heard earlier, works on a lot of very sensitive data and very sensitive workflows. And I think oftentimes something that I hear from engineers in the space is this concept of hallucination and safety, and how are you actually going to be able to produce a lot of these things to have benefits to downstream finance teams? And we're pretty big believers that it all comes down to the catalog of tools that teams are building and integrating with on a daily basis.
And so, what you're seeing here is our internal tool catalog. So, an example would be get a policy snippet or per diem rate or recent transactions. And these are built alongside product teams to really understand a lot of the nuances in the data and the use case. And what's really cool about this is not only can you see where there's gaps in our offering, that, oh, we actually don't have a tool for this specific use case. These can be used both in internal repos and our core product. And so, if you have an idea of, I want to do a cool reimbursement agent idea, here are the different ways to integrate the tools, the different APIs and systems that they integrate with, and now you can prototype that on a totally new product and vibe-coded surface area without having to worry about learning all of these things from scratch or building the tools on your own. We're up to many hundreds of these tools today, and, as Nik mentioned earlier, we think that this could be multiple thousands over time.
On the topic of context, another big thing we think about is context for our customers: how do we actually integrate the financial stack and allow them to be a little more productive. We noticed a very similar problem internally on our engineering team. And I think something that's not always obvious is that, even if you're using something like Claude Code or Codex, there's all this fragmentation of what you actually do on a daily basis to get work done in your company that's not integrated, too. There's logs in Datadog, there's a production database that has a bunch of things going on, there's different alerting systems, there's incident.io, there's a Slack message you have to pull in, there's a Notion doc, and then there's a lot of knowledge that those specific product teams have of how they actually need to get work done as well.
And so, at the end of last year, we decided to try to solve this problem of how can we actually integrate all this context and build our own internal background coding agent, which we've called Ramp Inspect. You may have seen this on LinkedIn or X. We actually have open-sourced the blueprint of how we built this, and at the end I can definitely show you guys a link of where to find that. And the progress has been pretty phenomenal of actually integrating this into a background agent that can run autonomously as people are in meetings, as bug fixes come up, and things like that. And currently this month, Ramp Inspect is responsible for over 50% of PRs that we merge to production.
I have some interesting, we're really big nerds with stats and numbers and things like that, so we have this dashboard to create one, subtle healthy competition, but also inspire people that they can actually use this as well. And as you can see, engineering has a huge lead in the amount of sessions, but you also have product, you also have design, there's risk, legal, corporate finance, and even marketing and CX teams using Ramp Inspect. And they're doing things like simple copy changes, they're doing logic fixes, they're trying to respond to incidents or bugs.
And what's been really cool to see as this has evolved over time, whoop, is how we've actually designed a couple of these things with some core principles to be really powerful. So what you're seeing here is a Ramp Inspect session. I think this is an example of a query that we were trying to fix. This spins up in the background a really fast Modal code sandbox. This allows us to resume, spin up, and spin down these containers in an isolated environment which has the same environment that you would have if you're developing at Ramp. There's a series of tasks to keep it on track, and it creates a GitHub branch and integrates with all of the context documents, our Datadog, our read replica so we can actually write queries, and different context documents that product teams have put together.
And what's really, I think, subtle about how we've designed this is we've designed it to be multiplayer first. And that means that as you integrate or you try to pair with a designer or somebody on the PM team, you can actually help them level up their own prompting skills. They can give us feedback of, hey, click on this link, this actually failed in a way that I wasn't expecting. And so that can be a really great source of cross-functional collaboration. That was a very subtle design choice that we made that ended up being a really big impact for the company.
And then these can be kicked off either via a Kanban UI, we have an API, and then also a Slack thread. And we can take the full context of the Slack thread when it is actually kicked off, so you don't have to re-prompt it with a bunch of conversation that happened earlier.
What you see here is we also have a full VS Code environment. We run VNC inside of a Modal sandbox as well, so this allows us to have Chrome DevTools and MCP so we can actually do full stack work, which is pretty cool. And it has access to the 150 plus thousand tests that we have. So, it also knows if things are broken, can respond to the CI inside of GitHub, and actually patch fixes before it actually pings you that the PR is done.
The link for this is builders.ramp.com. I think it's like one of the first blog posts that we have, or the most recent blog post that we have, and we open source like the whole blueprint of how to build this and put this together as well. I think there's also a GitHub repo called Open Inspect, which is an open source implementation of this as well.
So, it's been pretty interesting to see the impact that Ramp Inspect has had. We're over 50% of PRs that we merge on a weekly basis goes through the system. And so, with all this time not spent on thinking about these really low-level firefighting tasks or really low-level small fixes or tweaks that can be kind of democratized across the company, we're really rethinking like how our engineering teams operate and think about their job and how they can actually be really impactful in this new kind of AI-native future.
And so, as a thought experiment, let's pretend we have two different teams. I'm sure everyone in this room has worked with like their handful of extraordinary teams, maybe teams that are finding their footing. And you'll notice that there's like a couple of different qualities that may resonate.
So, we have team A on the left here. And let's say that they really care about impact, they handle ambiguous problems, they understand the product, business, and data, they adopt new tools, they can find creative solutions, and they obsess over like the user experience.
And then team B may also resonate with some people. You know, they debate libraries, they add process when things start to feel chaotic. They constantly complain about head count. They bike shed the details instead of actually focusing on the user experience. Like, hey, should we use, you know, functional programming paradigm here or what version of, you know, different TypeScript libraries do we want to use? And then they build before understanding the... right? They just say, "Hey, we're going to just vibe code this, bro. Don't worry." Or they focus on, you know, performative code quality or nitpicks that may very much be a subjective kind of matter of fact, as well. I've worked on both of these teams, and I think the argument that I'm going to make today is that there's going to be a divergence, I think, depending on what side of the aisle you land there.
This is a study from Harvard that was out, I think, the end of last year. And it was very much geared towards juniors and seniors in terms of what's actually happening with hiring trends in engineering since AI tools have accelerated. And I think what this glosses over is that I don't think it's just a years of experience problem. I actually think it's very much all of the different qualities that I said in team A versus team B that really make it apparent that coding was never really the hardest part of a lot of jobs for a long time. There's all these other engineering principles that become really important than just raw coding speed.
So, when you think about like a staff or a staff plus engineer, you're really compensating those people more for a lot of the judgment that they bring to the table, the context, the ability to see around corners, all the learning that they have, the actual like scar tissue. And so, if, you know, you ask Opus 4.6 to do something, they'll have the knowledge to actually know if that is not going to work or that's actually a bad idea. And I think one thing that a lot of the narratives that we see in the media get wrong about coding agents is they don't really identify the fact that you could still build the wrong thing just a lot faster, and you can build like bigger messes. And I think that having a lot of these skills of a team A and really focusing on like what is the context and reason behind this will only become more important in AI.
And so, what does that actually look like? We hit on some of these things. Figuring out what to build and understanding users well enough. Selling an idea to skeptical stakeholders. This is still something... when we decided to build a background coding agent, this was not something that was obvious that we should be spending time on. Having good design decisions with incomplete information and maintaining momentum through the long middle of this project, which can be really gnarly.
And I think this last bit, you know, everyone in this room I'm sure is painfully aware of, you know, the conversation around SaaS and the stock market and things like that. And I think this is like a big element that they gloss over, which is that yes, it's easy to vibe code something, but actually going through that middle process is like why you need really good engineers to actually get something deployed that has product market fit, that people are really excited about. And I think not enough people recognize that.
And so where does that leave us? Personally, I think there's a lot of kind of doomerism and scariness around a lot of the AI narratives, but I think it's also a really exciting time to be building. Unlike maybe factory work or farming, software's never done. We have this really kind of like meme internally where we say, you know, job's not finished. You've probably seen it in the marketing as well. I think software's perpetually not finished. And so with all this extra capacity, with people focusing less on this kind of low-level work and more on high-leverage engineering tasks, I think four things are going to really happen.
I think companies are just going to chase opportunities they couldn't afford to pursue. I don't know if we would be chasing these like agentic workflows and really thinking about bigger scale problems in the financial stack if this technology didn't exist. People are going to enter adjacent markets. They're going to try to stitch together more value for customers. It's not going to be like because everyone's 2x more productive, you need half the people.
You're going to rebuild systems that are too expensive to touch. I think building an internal background coding agent for a company that does financial operations software felt like probably a pretty crazy idea, but now that makes a ton of sense.
And raise the bar for what good enough means. I think, you know, being able to kind of build more mind-blowing experiences for users, provide a lot more value is going to be the narrative of the next decade. And I'm super excited to be able to build some of these things and see what everyone in this room is going to build, too. So, thank you.
Article published
