From Napkin Math to turbopuffer: Simon Eskildsen on First Principles, Simple Software, and Honest Fundraising

Open on YouTube ↗
Overview

At AI Engineer's World Fair in San Francisco, Gergely Orosz interviewed Simon Eskildsen, co-founder and CEO of turbopuffer, a search and vector database built on object storage. The conversation covered Eskildsen's path from teenage hobbyist to eight years on Shopify's infrastructure team, the "napkin math" habit that shaped how he evaluates systems, and how turbopuffer went from a single-server side project to a company whose first customer was Cursor. Eskildsen's consistent position was that you should reason from what the hardware can theoretically do, build the simplest thing that works, and be explicit, even with investors, about why you are doing what you are doing.

27 min read

From PowerPoint to the Informatics Olympiad

Asked where he fell in love with computers, Eskildsen answered: PowerPoint. Slides that jump to other slides when you click a shape become "Turing complete real quick," he said, and he used that to build convoluted games. From there he found FrontPage in the Office suite. He remembered the heartbreak of someone opening one of his sites in Firefox and seeing it fall apart, since it only rendered properly in Internet Explorer. One day he accidentally opened FrontPage's HTML view, saw code he couldn't understand, and started searching online for snippets, such as ones that changed the cursor. That led to Dreamweaver, then to PHP once he wanted dynamic pages.

At around 11 or 12, he said, he "exhausted the internet on Danish-language programming advice." He then spent four years hooked on World of Warcraft, which he credited with making him very good at English and opening up the far larger English-language web. He mused that today an LLM would simply explain things to him in Danish and he would never have hit that wall.

In high school he took on programming jobs and worked for a startup. An internet friend on the Australian team told him about the International Olympiad in Informatics (IOI) and suggested Denmark probably had a team too. He found the application on "some little mysterious website" and began solving algorithmic problems that looked nothing like his HTML and PHP work. To illustrate the style, he described a problem that isn't an actual IOI task: given N trucks and M packages with various dimensions, decide which packages go in which truck as well as possible. It is NP-complete and can't be solved exactly, but contestants compete to produce the best answer.

Getting Hired by Shopify Through a Nokia Article

Shopify found him while he was still in high school, and not through open source. In 2013 he dropped his iPhone, which back then usually meant the screen was finished, and went back to an old Nokia brick phone. Before people widely talked about the downsides of smartphones, he wrote an article about the experience: he was calling people again and had his sense of direction back. It appeared briefly on Hacker News, and the New York Times featured it. A Shopify recruiter noticed the traffic and got in touch.

He doubts they realized he was still in high school. They invited him to Ottawa; he recalled his email reply asking something like "What's an Ottawa?" Walking into the building "just felt right." He told them he had to finish high school first, then moved to Canada in 2013.

He thought it would be a gap year before going back to study. But he felt very insecure about not having a computer science degree, with the IOI as his only formal exposure. He said the olympiad had been a good crash course, and above all it had taught him that you can sit down with a paper and figure it out if you spend enough time. So during his first year at Shopify, he wrote down every term he didn't know and read about it that evening. He assumed that anyone who mentioned TCP must know the three-way handshake, how TLS layers on top, and how it all looks in Wireshark. He doesn't think that was actually true, but it pushed him to learn it all. He soon realized he had already found what he wanted to do and saw no reason to leave and come back later.

He now looks for the same trait when interviewing engineers: someone who "can't help themselves" from peeling back layers. For him that led to infrastructure. Although he worked on the product side, he sat with the infrastructure engineers at lunch to learn what they were talking about. He joked that he still can't explain what is "reverse" about a reverse proxy, or what is "inverted" about an inverted index, and called them terrible names.

Scaling Shopify's Infrastructure

Eskildsen said he felt fortunate to have a front-row seat to the rapid scaling of 2010s SaaS companies. He joined the infrastructure team around 2013–2014, as Docker arrived and Shopify was containerizing everything. The company was growing 120–140% a year, a rate he noted can look quaint next to today's AI companies. Every Black Friday was expected to be much harder than the one before. Shopify was still buying physical hardware then, so orders had to be placed at a specific time based on growth estimates.

In his experience, application-layer scaling problems usually come back to the database. He ended up working in the layer between Rails and the databases. Shopify at the time mostly orchestrated its databases rather than patching them. He quoted his boss Camilo: "You can't cache writes." At some point you have to go beyond a single shard. Shopify sharded around when he joined, and he recalled the cutover happening about a week before Black Friday, which he called mind-blowing, but it worked. Later projects included running in multiple data centers.

He also described a mysterious Redis server with 128 GB of RAM, a lot at the time, which nobody fully understood because people had treated it as a general key-value store. When it went down one day, the team found this terrifying and began splitting it apart.

Resilience Testing and Toxiproxy

That incident fed into a broader effort. If the component that stores sessions for a Shopify storefront goes down, the whole store shouldn't go down with it. But that is the default failure mode, Eskildsen said, unless the language forces you to decide how to handle each failure. The team built a matrix of how each service should behave when each component is unavailable, and he ended up writing the test suite for much of it.

He didn't want to rely on mocks. His first idea was to shell out to GDB, attach to the running process, and close the file descriptor to the database, simulating a database failure through the whole stack. He admitted it was "a little crazy" and never ran in CI. Still, it uncovered many failure-handling problems at the connection layer, which led to fixes upstreamed to Rails and similar projects.

The next step was Toxiproxy. He described it as a layer-4 proxy between the application and its databases, correcting himself after first calling it layer 7. It passes traffic through without needing to understand the MySQL protocol, but exposes an API to take the database down or make it slow. Later it gained more layer-7-style failure modes. Because the real drivers still talk to a real connection, the tests exercise the drivers' own failure handling instead of mocking it. That made the whole resilience matrix testable in CI. A test could say, in effect, "with MySQL's sessions table down, load this page or run a checkout," passing a lambda with the actions to perform.

According to Eskildsen, this uncovered dozens of issues in the MySQL driver and Rails. Nobody in the ecosystem had been testing these paths, and they're hard to study in production because, during a real outage, everyone is focused on recovering rather than asking what the application should have done. Orosz pointed out that the hardest problems in large systems usually involve state, and state is hard to test before it breaks. Eskildsen said that to his knowledge Toxiproxy still runs in Shopify's CI.

Leaving After Eight Years

Eskildsen stayed at Shopify from 2013 to 2021. He had been there since he was 18, and apart from one high-school startup it was the only company he knew. He wanted to learn faster and felt it was time to "inject some novelty."

By then he had worked on caching, multi-data-center operation, and many database scaling projects. With Justine, now his co-founder, he rewrote Shopify's storefront, which he said served almost 100% of traffic 18 months after they began. He joked that much of the scaling pressure came from the Kardashians launching products on Shopify.

The Napkin Math Project

Before leaving, and more intensively afterwards, Eskildsen maintained a project he called napkin math: a table on GitHub, generated by a Rust script, of about 50 numbers. It included how much bandwidth you can drive to DRAM, an NVMe SSD, or an EBS volume, and how long an S3 round trip takes and what it costs. It also covered prices. He cited roughly $2 per GB for memory, 2 cents per GB for S3, and 10 cents per GB for disk, along with spot and three-year-commit pricing. He made flashcards for nearly every cell so he would know the numbers by heart.

He started it because of his role reviewing infrastructure plans for product teams. Teams would often say they had benchmarked database A, found it poor, and chosen database B. "I hate benchmarks so much," he said, because a benchmark result alone didn't satisfy him when it conflicted with his intuition. His example was a search query: three terms, a known number of matching documents per term, so a known number of megabytes of posting lists to intersect. With about 100 GB/s of DRAM bandwidth across cores, that should take around 10 ms. If the benchmark shows 10 seconds, "one of us is wrong." Either his understanding had a gap, which he said was very likely, or the wrong thing had been benchmarked. For instance, the benchmark might be running a distributed query across 100 nodes, which naturally makes the P99 very high.

In those reviews he wanted "ammo" to do the calculation on the spot. For example: how a B-tree works, how many pages a lookup visits, roughly 1 ms per random SSD read, how many reads that adds up to. Then the question becomes what explains the gap between that estimate and the observed query time. Is the query plan wrong? Is there a MySQL bug? Are the disks bad?

The fsync Puzzle

After leaving Shopify he wrote many articles in this style. One question he described in detail: shouldn't MySQL's maximum writes per second equal the number of fsyncs per second, since each write must be persisted? If an fsync takes about 1 ms, that suggests about 1,000 writes per second, which seemed too low. Testing on a small machine, he measured about 10,000 writes per second.

The answer is batching. He spent roughly 24 obsessive hours on it, before LLMs, writing BPF traces, and found that each fsync was much larger than he had expected. He then read the MySQL code and found an obscure article on the relevant internals by someone in a small German town. He joked that he is convinced "the entire internet runs on small towns in Bavaria."

Three Ingredients Behind turbopuffer

Eskildsen said three things came together to lead to turbopuffer.

The first was his final Shopify project, search, which he "didn't have a good time" with. Without naming the vendor, he described a traditional search engine that was hard to operate and that he couldn't get to perform anywhere near his napkin-math estimates. It had no query planner. Sometimes performance tracked expectations and sometimes it didn't at all, even after he read its source code. He never expected to work on search again.

The second was the napkin math project itself, which gave him a feel for what a machine can achieve when used properly.

The third came after Shopify, through what he called "angel engineering": instead of investing money in friends' companies, he worked with them and vested equity, because he wanted to be hands-on and see other environments. The same problem kept coming up. After ChatGPT launched in 2022, he worked with Readwise, a bootstrapped Canadian company whose product lets users save articles and search them later. They wanted to connect documents to AI, and with context windows of only 4–8 KB, fast search was essential.

He built a recommendation engine for them that worked surprisingly well. Running it on one co-founder's feed, with permission, he said, he learned from the recommendations that the co-founder's wife was pregnant. But when he estimated the cost of running it for all users, it came to about $30,000 a month, while Readwise spent roughly $5,000 a month on all other infrastructure combined. After adding the gross margin a company needs, it made no sense, so they didn't ship it. He moved on to tasks like tuning Postgres autovacuum.

He kept wondering why storing those vectors was so expensive. One day he did the napkin math on putting everything in S3, clustering the vectors, and organizing the files carefully, and concluded it might work.

Summer 2023: The Simplest Possible Version on S3

He spent the summer of 2023 "hammering my head against the wall" to reach acceptable latency. S3 is very durable but slow. He cited a P99 of about 200 ms for a 256 or 512 KB object. He stressed designing for P99, or even P99.9, because a query on S3 usually involves many requests. If you walk a tree stored on S3, each level costs another ~200 ms, so the design has to minimize round trips.

In July 2023 he had something working end to end, rewrote it about twice, and released it in October 2023. At that point it was a project, not a company. He said he "barely knew what a VC was." It was driven by curiosity and a feeling that if he didn't build it, someone else would.

He described himself as very pragmatic, and the first version was the simplest thing he could build. He skimmed the LSM-tree literature for the basic idea and barely implemented it. The design: run a clustering algorithm on the vectors, write each cluster to its own file ("cluster one, cluster two, cluster three"), and write a separate file of centroids. A search downloads the centroids, finds the nearest ones, and downloads the N closest clusters. There were some optimizations, such as merging adjacent clusters into files to control cost and performance, but that was essentially it.

He didn't write a caching layer at first. He put an Nginx reverse proxy in front of S3 to cache objects, since he knew Nginx well and had written a lot of Nginx Lua. To delete items from the cache, he shelled out to xargs and removed files after reverse-engineering Nginx's cache directory layout. The whole thing ran on a single server inside a tmux session. His attitude was: let's see if anyone cares.

How Cursor Became the First Customer

Orosz explained that he first heard of turbopuffer while researching Cursor's backend for a deep dive. Cursor described trying Postgres, then Amazon Aurora, which surprisingly didn't work well for them, before moving to turbopuffer, and mentioned being one of its first customers. Eskildsen corrected that: Cursor was the first customer.

He announced turbopuffer on Twitter while fed up after a summer of work. He only wanted to continue if anyone cared. It was still a single tmux instance on an eight-core node in GCP, and he planned to set it up properly only if someone went to production. He called it "the MVP of MVP" and said anyone who had worked on database internals would have had too much pride to ship it. But he treated it like a SaaS project and saw no reason a database couldn't be built that way. He knew how to run highly available software once there was a reason to. His pitch was a million vectors for a dollar, when he believed the cheapest working alternative was about $100 per million. Despite the simplicity, he said the durability guarantees were already the same as today: writes were committed directly, and shutting down every VM would lose no data.

Cursor reached out. Eskildsen guessed it had about eight people then. Knowing the founders now, he speculated that they had already noticed their unit economics, with all vectors held in DRAM, didn't work. They had likely wondered why nobody had built a system that keeps actively used codebases in memory, leaves the rest in object storage, and loads them into cache on demand, so that a codebase is in RAM a few seconds after you open it. He noted that one co-founder had tweeted early on about using S3 for caching, which he said still few people do despite the favorable economics, and he thought they might have been considering building it themselves. He stressed that this was his guess, not something he knew.

After exchanging emails, and knowing nothing about B2B sales at the time ("Now, I love B2B sales"), he flew from Canada to San Francisco and showed up at Cursor's office. They happened to be discussing a Postgres problem. He asked if they used pganalyze. They didn't, so they set it up. He said it was "the same thing as it always is with Postgres": autovacuum hadn't run enough, so queries were going to the heap instead of using index scans. He believes helping them with their database built enough trust that they figured he could also build one.

By then he had recruited Justine, whom he called the best engineer who ever worked at Shopify. Her first change was replacing the Nginx reverse-proxy cache with a direct file-based cache. That night Cursor decided to migrate and moved everything over a week or two. Eskildsen had promised to cut their bill by 95%, and he said the first turbopuffer bill was 95% lower than the last bill from their previous vendor. Orosz added that, according to his own conversations with Cursor, reliability had been the main pain point with the previous setup.

Orosz shared a comment from a Cursor co-founder: among the things you should never do is bet your business on a tiny startup where you are the only or biggest customer, "except for turbopuffer." Orosz's takeaway was that high-quality work can open doors, and that startups can sometimes take seemingly irrational risks when they have conviction. He suggested Eskildsen created that conviction by showing up in person and demonstrating what he knew.

Asking Jensen Huang If He Vapes

turbopuffer runs mainly on CPUs. Eskildsen described presenting at an NVIDIA event where a few companies talked about their businesses and possible partnerships to Jensen Huang and a group of NVIDIA leadership. Nervous and in "a goofy mood," he opened by saying that if everything went south, turbopuffer could always pivot into vapes. Huang replied, "Judging by your slide, maybe you should." Not knowing what else to say, Eskildsen asked, "Well, Jensen, do you vape?" Huang didn't answer. Someone on the team then messaged the whole company that Simon had just asked Jensen if he vapes.

His team had warned him beforehand not to say "the C-word," CPUs. He then "couldn't stop talking about CPUs": how great AVX-512 is, how much they love SIMD, how easy CPUs are to get. He said he stopped just short of saying he was glad not to need GPUs. According to Eskildsen, Huang took an interest in that.

Why CPUs Are Suddenly Scarce

Orosz assumed that while GPUs are scarce, CPUs should be easy to get. Eskildsen said, "No. It's not anymore." He declined to speculate much about the macro picture but offered his explanation. Reinforcement learning is becoming a large share of AI workloads, and it needs many CPUs: to teach a model to search, use grep, or start a bash shell, it has to run real programs and learn from the results. Agents also run on CPUs because they do general-purpose work. As AI is applied to more domains, gaps appear. He gave the hypothetical examples of models not being good at CAD or shipbuilding, which then requires even more RL environments. NVMe SSDs are also needed, and DRAM is tight partly because GPU servers need a lot of it.

He expects the CPU situation to get "a lot worse before it gets a lot better." He said even large companies compete with each other for allocations, and turbopuffer sometimes competes for CPUs with the same companies it sells to. Orosz mentioned a customer dinner the night before where Reflection said it had maxed out its purchases on the longest possible contracts. Eskildsen said turbopuffer works with the cloud providers on which regions have CPU capacity, which comes down to where power is available, and coordinates this with its largest customers.

He considers turbopuffer fortunate: it needs only some CPUs, NVMe SSDs, and S3. It could make architectural changes to reduce its exposure but would rather spend engineering effort elsewhere. Its main strength is running on many different SKUs rather than depending on one instance type. His current favorites are GCP's C4 instances, the Z4D instances (which he said perform very well after some optimization work), and the Arm-based C4A. He recalled from Shopify that ahead of Black Friday/Cyber Monday you had to tell cloud providers months in advance how much you would use: "The clouds are not infinite as they seem when you're small."

Eskildsen's View of Venture Capital

Orosz noted that turbopuffer had barely announced funding and asked how Eskildsen thinks about raising. Eskildsen went back to the start. He had promised Cursor that he and Justine could bring its bill down to about $4,000 a month. That figure came from rough napkin math about what a better version of turbopuffer should cost, and it became both the launch price and the guarantee to Cursor. The software was reliable but simple, and he said "simplicity above everything" is his core engineering principle. He and Orosz agreed that long tenures, his at Shopify and Orosz's at Uber, show that simplicity almost always wins as software ages.

At the time he wasn't sure this was a venture-scale opportunity. He understood that taking VC money means everyone expects a large return on some timeline, no matter how friendly the meeting, because investors answer to others, such as Canadian pension funds. The product felt niche, and he didn't know if it could be a billion-dollar company. That was fine with him. As a "dumb Danish person," he simply compared Cursor's bill with his GCP bill and wanted the first number to be larger. The plan was to optimize until the numbers roughly matched, find more workloads, and eventually pay himself and Justine. He also had no investor connections. He called himself an "outsider squared": from Aarhus, Denmark, an outsider in Canada, and from Canada an outsider to San Francisco. So he reasoned from first principles about what investors need and whether he could deliver it.

The trigger for raising was Boyan, whom Eskildsen knew from the IOI in 2012 and 2013. Boyan competed for North Macedonia and was so strong that his teammates called him "God." Eskildsen wanted to hire him but couldn't afford it. By then he and Justine had gone about six months without salaries and had spent tens of thousands of dollars on GCP. He called Locky, the one person he knew in Silicon Valley, and proposed raising about $700,000 in January. That would fund two engineers for the rest of the year, while he and Justine still went unpaid, plus some buffer. His terms were explicit: if there was no product-market fit and no sign of a big opportunity by year's end, they would shut down, take nothing, and return the money. He thinks it was the first time Locky had heard it put that way. Other VCs he told found it alarming, and he suspects that on the West Coast it sounded like low ambition. His explanation: "When I don't know how to play a game, I just play with open cards."

By then they were developing conviction that it could become very big, and they also didn't want to keep going unless it could. They hired Boyan and Morgan as the first engineers, and Eskildsen said the company became profitable later that year.

Six Reasons to Raise Capital

Eskildsen listed six reasons a company raises money and argued founders should be honest about which one applies:

  1. Funding R&D. This was turbopuffer's first raise. The founders had been funding R&D themselves through unpaid work and paying bills, and wanted to learn faster by hiring.
  2. Funding growth. You have built something and want to spend money telling the world about it.
  3. The founders' ego. He called this "very popular" and very dangerous: big numbers and press coverage dilute every employee and set the price that future employees' upside depends on. It can become a status game, "and that's not what it's about."
  4. Rewarding employees. On a long journey with the best people, whom he says are by definition few, you want to reward them. This was the reason for turbopuffer's December raise: letting employees sell some equity instead of waiting for a distant IPO or other event.
  5. Strategic partnerships. He said some partnerships formed in San Francisco have made companies.
  6. M&A or similar.

turbopuffer's two raises, he said, fell under reasons one and four. Orosz agreed that the ego and identity side is rarely discussed and becomes harder to avoid the closer you are to tech hubs where everyone is raising.

Running Fully Remote: Campfires and Turbocredits

Orosz noted that many AI companies prefer to be in person, often in San Francisco, for faster iteration, while turbopuffer has been fully remote from the start. Eskildsen said founding in 2023, soon after COVID, made remote work normal, and Shopify's infrastructure team had always been remote because it was hard to get everyone to move to Ottawa. In his view there are perhaps two cities where you can quickly build a database company, San Francisco and maybe New York. If you don't want to be there, you have to commit fully to a distributed model.

For turbopuffer, distributed doesn't mean never meeting. Everyone gathers twice a year, recently in Banff and Mexico City. Its more distinctive practice is the "campfire": when a few people happen to be in the same place, they declare a campfire and invite anyone who wants to come. The week of the talk was a campfire in San Francisco, with customer meetings and dinners. Attendance is optional. Some people want to stay home, focus, and see their families, and only attend the two offsites, which Eskildsen said fits the model perfectly. Others fly roughly every two weeks. He told of an employee in Ottawa who saw colleagues dialing in from a New York meeting room during a campfire, got such FOMO that she took an Uber to the airport and flew to New York to join them.

The company also created "turbocredits." Extracurricular contributions like a conference talk or blog post earn one, and a turbocredit upgrades your next flight to business class, which Eskildsen said further encourages time together. Engineers who choose to spend two days on a conference expo floor talking to customers, which he called taxing, also earn one. He said the credits may take on a life of their own: someone has already suggested a central bank, interest rates, and a betting market for them.

Closing

Orosz closed by observing that although many AI companies use turbopuffer as infrastructure, the conversation had been mostly about engineering principles, curiosity, and how people come to trust each other and work together, rather than about AI.