From Napkin Math to turbopuffer: Simon Eskildsen on First Principles, Simple Software, and Honest Fundraising
The Pragmatic EngineerAt AI Engineer's World Fair in San Francisco, Gergely Orosz interviewed Simon Eskildsen, co-founder and CEO of turbopuffer, a search and vector database built on object storage. The conversation covered Eskildsen's path from teenage hobbyist to eight years on Shopify's infrastructure team, the "napkin math" habit that shaped how he evaluates systems, and how turbopuffer went from a single-server side project to a company whose first customer was Cursor. Eskildsen's consistent position was that you should reason from what the hardware can theoretically do, build the simplest thing that works, and be explicit, even with investors, about why you are doing what you are doing.
From PowerPoint to the Informatics Olympiad
Asked where he fell in love with computers, Eskildsen answered: PowerPoint. Slides that jump to other slides when you click a shape become "Turing complete real quick," he said, and he used that to build convoluted games. From there he found FrontPage in the Office suite. He remembered the heartbreak of someone opening one of his sites in Firefox and seeing it fall apart, since it only rendered properly in Internet Explorer. One day he accidentally opened FrontPage's HTML view, saw code he couldn't understand, and started searching online for snippets, such as ones that changed the cursor. That led to Dreamweaver, then to PHP once he wanted dynamic pages.
At around 11 or 12, he said, he "exhausted the internet on Danish-language programming advice." He then spent four years hooked on World of Warcraft, which he credited with making him very good at English and opening up the far larger English-language web. He mused that today an LLM would simply explain things to him in Danish and he would never have hit that wall.
In high school he took on programming jobs and worked for a startup. An internet friend on the Australian team told him about the International Olympiad in Informatics (IOI) and suggested Denmark probably had a team too. He found the application on "some little mysterious website" and began solving algorithmic problems that looked nothing like his HTML and PHP work. To illustrate the style, he described a problem that isn't an actual IOI task: given N trucks and M packages with various dimensions, decide which packages go in which truck as well as possible. It is NP-complete and can't be solved exactly, but contestants compete to produce the best answer.
Getting Hired by Shopify Through a Nokia Article
Shopify found him while he was still in high school, and not through open source. In 2013 he dropped his iPhone, which back then usually meant the screen was finished, and went back to an old Nokia brick phone. Before people widely talked about the downsides of smartphones, he wrote an article about the experience: he was calling people again and had his sense of direction back. It appeared briefly on Hacker News, and the New York Times featured it. A Shopify recruiter noticed the traffic and got in touch.
He doubts they realized he was still in high school. They invited him to Ottawa; he recalled his email reply asking something like "What's an Ottawa?" Walking into the building "just felt right." He told them he had to finish high school first, then moved to Canada in 2013.
He thought it would be a gap year before going back to study. But he felt very insecure about not having a computer science degree, with the IOI as his only formal exposure. He said the olympiad had been a good crash course, and above all it had taught him that you can sit down with a paper and figure it out if you spend enough time. So during his first year at Shopify, he wrote down every term he didn't know and read about it that evening. He assumed that anyone who mentioned TCP must know the three-way handshake, how TLS layers on top, and how it all looks in Wireshark. He doesn't think that was actually true, but it pushed him to learn it all. He soon realized he had already found what he wanted to do and saw no reason to leave and come back later.
He now looks for the same trait when interviewing engineers: someone who "can't help themselves" from peeling back layers. For him that led to infrastructure. Although he worked on the product side, he sat with the infrastructure engineers at lunch to learn what they were talking about. He joked that he still can't explain what is "reverse" about a reverse proxy, or what is "inverted" about an inverted index, and called them terrible names.
Scaling Shopify's Infrastructure
Eskildsen said he felt fortunate to have a front-row seat to the rapid scaling of 2010s SaaS companies. He joined the infrastructure team around 2013–2014, as Docker arrived and Shopify was containerizing everything. The company was growing 120–140% a year, a rate he noted can look quaint next to today's AI companies. Every Black Friday was expected to be much harder than the one before. Shopify was still buying physical hardware then, so orders had to be placed at a specific time based on growth estimates.
In his experience, application-layer scaling problems usually come back to the database. He ended up working in the layer between Rails and the databases. Shopify at the time mostly orchestrated its databases rather than patching them. He quoted his boss Camilo: "You can't cache writes." At some point you have to go beyond a single shard. Shopify sharded around when he joined, and he recalled the cutover happening about a week before Black Friday, which he called mind-blowing, but it worked. Later projects included running in multiple data centers.
He also described a mysterious Redis server with 128 GB of RAM, a lot at the time, which nobody fully understood because people had treated it as a general key-value store. When it went down one day, the team found this terrifying and began splitting it apart.
Resilience Testing and Toxiproxy
That incident fed into a broader effort. If the component that stores sessions for a Shopify storefront goes down, the whole store shouldn't go down with it. But that is the default failure mode, Eskildsen said, unless the language forces you to decide how to handle each failure. The team built a matrix of how each service should behave when each component is unavailable, and he ended up writing the test suite for much of it.
He didn't want to rely on mocks. His first idea was to shell out to GDB, attach to the running process, and close the file descriptor to the database, simulating a database failure through the whole stack. He admitted it was "a little crazy" and never ran in CI. Still, it uncovered many failure-handling problems at the connection layer, which led to fixes upstreamed to Rails and similar projects.
The next step was Toxiproxy. He described it as a layer-4 proxy between the application and its databases, correcting himself after first calling it layer 7. It passes traffic through without needing to understand the MySQL protocol, but exposes an API to take the database down or make it slow. Later it gained more layer-7-style failure modes. Because the real drivers still talk to a real connection, the tests exercise the drivers' own failure handling instead of mocking it. That made the whole resilience matrix testable in CI. A test could say, in effect, "with MySQL's sessions table down, load this page or run a checkout," passing a lambda with the actions to perform.
According to Eskildsen, this uncovered dozens of issues in the MySQL driver and Rails. Nobody in the ecosystem had been testing these paths, and they're hard to study in production because, during a real outage, everyone is focused on recovering rather than asking what the application should have done. Orosz pointed out that the hardest problems in large systems usually involve state, and state is hard to test before it breaks. Eskildsen said that to his knowledge Toxiproxy still runs in Shopify's CI.
Leaving After Eight Years
Eskildsen stayed at Shopify from 2013 to 2021. He had been there since he was 18, and apart from one high-school startup it was the only company he knew. He wanted to learn faster and felt it was time to "inject some novelty."
By then he had worked on caching, multi-data-center operation, and many database scaling projects. With Justine, now his co-founder, he rewrote Shopify's storefront, which he said served almost 100% of traffic 18 months after they began. He joked that much of the scaling pressure came from the Kardashians launching products on Shopify.
The Napkin Math Project
Before leaving, and more intensively afterwards, Eskildsen maintained a project he called napkin math: a table on GitHub, generated by a Rust script, of about 50 numbers. It included how much bandwidth you can drive to DRAM, an NVMe SSD, or an EBS volume, and how long an S3 round trip takes and what it costs. It also covered prices. He cited roughly $2 per GB for memory, 2 cents per GB for S3, and 10 cents per GB for disk, along with spot and three-year-commit pricing. He made flashcards for nearly every cell so he would know the numbers by heart.
He started it because of his role reviewing infrastructure plans for product teams. Teams would often say they had benchmarked database A, found it poor, and chosen database B. "I hate benchmarks so much," he said, because a benchmark result alone didn't satisfy him when it conflicted with his intuition. His example was a search query: three terms, a known number of matching documents per term, so a known number of megabytes of posting lists to intersect. With about 100 GB/s of DRAM bandwidth across cores, that should take around 10 ms. If the benchmark shows 10 seconds, "one of us is wrong." Either his understanding had a gap, which he said was very likely, or the wrong thing had been benchmarked. For instance, the benchmark might be running a distributed query across 100 nodes, which naturally makes the P99 very high.
In those reviews he wanted "ammo" to do the calculation on the spot. For example: how a B-tree works, how many pages a lookup visits, roughly 1 ms per random SSD read, how many reads that adds up to. Then the question becomes what explains the gap between that estimate and the observed query time. Is the query plan wrong? Is there a MySQL bug? Are the disks bad?
The fsync Puzzle
After leaving Shopify he wrote many articles in this style. One question he described in detail: shouldn't MySQL's maximum writes per second equal the number of fsyncs per second, since each write must be persisted? If an fsync takes about 1 ms, that suggests about 1,000 writes per second, which seemed too low. Testing on a small machine, he measured about 10,000 writes per second.
The answer is batching. He spent roughly 24 obsessive hours on it, before LLMs, writing BPF traces, and found that each fsync was much larger than he had expected. He then read the MySQL code and found an obscure article on the relevant internals by someone in a small German town. He joked that he is convinced "the entire internet runs on small towns in Bavaria."
Three Ingredients Behind turbopuffer
Eskildsen said three things came together to lead to turbopuffer.
The first was his final Shopify project, search, which he "didn't have a good time" with. Without naming the vendor, he described a traditional search engine that was hard to operate and that he couldn't get to perform anywhere near his napkin-math estimates. It had no query planner. Sometimes performance tracked expectations and sometimes it didn't at all, even after he read its source code. He never expected to work on search again.
The second was the napkin math project itself, which gave him a feel for what a machine can achieve when used properly.
The third came after Shopify, through what he called "angel engineering": instead of investing money in friends' companies, he worked with them and vested equity, because he wanted to be hands-on and see other environments. The same problem kept coming up. After ChatGPT launched in 2022, he worked with Readwise, a bootstrapped Canadian company whose product lets users save articles and search them later. They wanted to connect documents to AI, and with context windows of only 4–8 KB, fast search was essential.
He built a recommendation engine for them that worked surprisingly well. Running it on one co-founder's feed, with permission, he said, he learned from the recommendations that the co-founder's wife was pregnant. But when he estimated the cost of running it for all users, it came to about $30,000 a month, while Readwise spent roughly $5,000 a month on all other infrastructure combined. After adding the gross margin a company needs, it made no sense, so they didn't ship it. He moved on to tasks like tuning Postgres autovacuum.
He kept wondering why storing those vectors was so expensive. One day he did the napkin math on putting everything in S3, clustering the vectors, and organizing the files carefully, and concluded it might work.
Summer 2023: The Simplest Possible Version on S3
He spent the summer of 2023 "hammering my head against the wall" to reach acceptable latency. S3 is very durable but slow. He cited a P99 of about 200 ms for a 256 or 512 KB object. He stressed designing for P99, or even P99.9, because a query on S3 usually involves many requests. If you walk a tree stored on S3, each level costs another ~200 ms, so the design has to minimize round trips.
In July 2023 he had something working end to end, rewrote it about twice, and released it in October 2023. At that point it was a project, not a company. He said he "barely knew what a VC was." It was driven by curiosity and a feeling that if he didn't build it, someone else would.
He described himself as very pragmatic, and the first version was the simplest thing he could build. He skimmed the LSM-tree literature for the basic idea and barely implemented it. The design: run a clustering algorithm on the vectors, write each cluster to its own file ("cluster one, cluster two, cluster three"), and write a separate file of centroids. A search downloads the centroids, finds the nearest ones, and downloads the N closest clusters. There were some optimizations, such as merging adjacent clusters into files to control cost and performance, but that was essentially it.
He didn't write a caching layer at first. He put an Nginx reverse proxy in front of S3 to cache objects, since he knew Nginx well and had written a lot of Nginx Lua. To delete items from the cache, he shelled out to xargs and removed files after reverse-engineering Nginx's cache directory layout. The whole thing ran on a single server inside a tmux session. His attitude was: let's see if anyone cares.
How Cursor Became the First Customer
Orosz explained that he first heard of turbopuffer while researching Cursor's backend for a deep dive. Cursor described trying Postgres, then Amazon Aurora, which surprisingly didn't work well for them, before moving to turbopuffer, and mentioned being one of its first customers. Eskildsen corrected that: Cursor was the first customer.
He announced turbopuffer on Twitter while fed up after a summer of work. He only wanted to continue if anyone cared. It was still a single tmux instance on an eight-core node in GCP, and he planned to set it up properly only if someone went to production. He called it "the MVP of MVP" and said anyone who had worked on database internals would have had too much pride to ship it. But he treated it like a SaaS project and saw no reason a database couldn't be built that way. He knew how to run highly available software once there was a reason to. His pitch was a million vectors for a dollar, when he believed the cheapest working alternative was about $100 per million. Despite the simplicity, he said the durability guarantees were already the same as today: writes were committed directly, and shutting down every VM would lose no data.
Cursor reached out. Eskildsen guessed it had about eight people then. Knowing the founders now, he speculated that they had already noticed their unit economics, with all vectors held in DRAM, didn't work. They had likely wondered why nobody had built a system that keeps actively used codebases in memory, leaves the rest in object storage, and loads them into cache on demand, so that a codebase is in RAM a few seconds after you open it. He noted that one co-founder had tweeted early on about using S3 for caching, which he said still few people do despite the favorable economics, and he thought they might have been considering building it themselves. He stressed that this was his guess, not something he knew.
After exchanging emails, and knowing nothing about B2B sales at the time ("Now, I love B2B sales"), he flew from Canada to San Francisco and showed up at Cursor's office. They happened to be discussing a Postgres problem. He asked if they used pganalyze. They didn't, so they set it up. He said it was "the same thing as it always is with Postgres": autovacuum hadn't run enough, so queries were going to the heap instead of using index scans. He believes helping them with their database built enough trust that they figured he could also build one.
By then he had recruited Justine, whom he called the best engineer who ever worked at Shopify. Her first change was replacing the Nginx reverse-proxy cache with a direct file-based cache. That night Cursor decided to migrate and moved everything over a week or two. Eskildsen had promised to cut their bill by 95%, and he said the first turbopuffer bill was 95% lower than the last bill from their previous vendor. Orosz added that, according to his own conversations with Cursor, reliability had been the main pain point with the previous setup.
Orosz shared a comment from a Cursor co-founder: among the things you should never do is bet your business on a tiny startup where you are the only or biggest customer, "except for turbopuffer." Orosz's takeaway was that high-quality work can open doors, and that startups can sometimes take seemingly irrational risks when they have conviction. He suggested Eskildsen created that conviction by showing up in person and demonstrating what he knew.
Asking Jensen Huang If He Vapes
turbopuffer runs mainly on CPUs. Eskildsen described presenting at an NVIDIA event where a few companies talked about their businesses and possible partnerships to Jensen Huang and a group of NVIDIA leadership. Nervous and in "a goofy mood," he opened by saying that if everything went south, turbopuffer could always pivot into vapes. Huang replied, "Judging by your slide, maybe you should." Not knowing what else to say, Eskildsen asked, "Well, Jensen, do you vape?" Huang didn't answer. Someone on the team then messaged the whole company that Simon had just asked Jensen if he vapes.
His team had warned him beforehand not to say "the C-word," CPUs. He then "couldn't stop talking about CPUs": how great AVX-512 is, how much they love SIMD, how easy CPUs are to get. He said he stopped just short of saying he was glad not to need GPUs. According to Eskildsen, Huang took an interest in that.
Why CPUs Are Suddenly Scarce
Orosz assumed that while GPUs are scarce, CPUs should be easy to get. Eskildsen said, "No. It's not anymore." He declined to speculate much about the macro picture but offered his explanation. Reinforcement learning is becoming a large share of AI workloads, and it needs many CPUs: to teach a model to search, use grep, or start a bash shell, it has to run real programs and learn from the results. Agents also run on CPUs because they do general-purpose work. As AI is applied to more domains, gaps appear. He gave the hypothetical examples of models not being good at CAD or shipbuilding, which then requires even more RL environments. NVMe SSDs are also needed, and DRAM is tight partly because GPU servers need a lot of it.
He expects the CPU situation to get "a lot worse before it gets a lot better." He said even large companies compete with each other for allocations, and turbopuffer sometimes competes for CPUs with the same companies it sells to. Orosz mentioned a customer dinner the night before where Reflection said it had maxed out its purchases on the longest possible contracts. Eskildsen said turbopuffer works with the cloud providers on which regions have CPU capacity, which comes down to where power is available, and coordinates this with its largest customers.
He considers turbopuffer fortunate: it needs only some CPUs, NVMe SSDs, and S3. It could make architectural changes to reduce its exposure but would rather spend engineering effort elsewhere. Its main strength is running on many different SKUs rather than depending on one instance type. His current favorites are GCP's C4 instances, the Z4D instances (which he said perform very well after some optimization work), and the Arm-based C4A. He recalled from Shopify that ahead of Black Friday/Cyber Monday you had to tell cloud providers months in advance how much you would use: "The clouds are not infinite as they seem when you're small."
Eskildsen's View of Venture Capital
Orosz noted that turbopuffer had barely announced funding and asked how Eskildsen thinks about raising. Eskildsen went back to the start. He had promised Cursor that he and Justine could bring its bill down to about $4,000 a month. That figure came from rough napkin math about what a better version of turbopuffer should cost, and it became both the launch price and the guarantee to Cursor. The software was reliable but simple, and he said "simplicity above everything" is his core engineering principle. He and Orosz agreed that long tenures, his at Shopify and Orosz's at Uber, show that simplicity almost always wins as software ages.
At the time he wasn't sure this was a venture-scale opportunity. He understood that taking VC money means everyone expects a large return on some timeline, no matter how friendly the meeting, because investors answer to others, such as Canadian pension funds. The product felt niche, and he didn't know if it could be a billion-dollar company. That was fine with him. As a "dumb Danish person," he simply compared Cursor's bill with his GCP bill and wanted the first number to be larger. The plan was to optimize until the numbers roughly matched, find more workloads, and eventually pay himself and Justine. He also had no investor connections. He called himself an "outsider squared": from Aarhus, Denmark, an outsider in Canada, and from Canada an outsider to San Francisco. So he reasoned from first principles about what investors need and whether he could deliver it.
The trigger for raising was Boyan, whom Eskildsen knew from the IOI in 2012 and 2013. Boyan competed for North Macedonia and was so strong that his teammates called him "God." Eskildsen wanted to hire him but couldn't afford it. By then he and Justine had gone about six months without salaries and had spent tens of thousands of dollars on GCP. He called Locky, the one person he knew in Silicon Valley, and proposed raising about $700,000 in January. That would fund two engineers for the rest of the year, while he and Justine still went unpaid, plus some buffer. His terms were explicit: if there was no product-market fit and no sign of a big opportunity by year's end, they would shut down, take nothing, and return the money. He thinks it was the first time Locky had heard it put that way. Other VCs he told found it alarming, and he suspects that on the West Coast it sounded like low ambition. His explanation: "When I don't know how to play a game, I just play with open cards."
By then they were developing conviction that it could become very big, and they also didn't want to keep going unless it could. They hired Boyan and Morgan as the first engineers, and Eskildsen said the company became profitable later that year.
Six Reasons to Raise Capital
Eskildsen listed six reasons a company raises money and argued founders should be honest about which one applies:
- Funding R&D. This was turbopuffer's first raise. The founders had been funding R&D themselves through unpaid work and paying bills, and wanted to learn faster by hiring.
- Funding growth. You have built something and want to spend money telling the world about it.
- The founders' ego. He called this "very popular" and very dangerous: big numbers and press coverage dilute every employee and set the price that future employees' upside depends on. It can become a status game, "and that's not what it's about."
- Rewarding employees. On a long journey with the best people, whom he says are by definition few, you want to reward them. This was the reason for turbopuffer's December raise: letting employees sell some equity instead of waiting for a distant IPO or other event.
- Strategic partnerships. He said some partnerships formed in San Francisco have made companies.
- M&A or similar.
turbopuffer's two raises, he said, fell under reasons one and four. Orosz agreed that the ego and identity side is rarely discussed and becomes harder to avoid the closer you are to tech hubs where everyone is raising.
Running Fully Remote: Campfires and Turbocredits
Orosz noted that many AI companies prefer to be in person, often in San Francisco, for faster iteration, while turbopuffer has been fully remote from the start. Eskildsen said founding in 2023, soon after COVID, made remote work normal, and Shopify's infrastructure team had always been remote because it was hard to get everyone to move to Ottawa. In his view there are perhaps two cities where you can quickly build a database company, San Francisco and maybe New York. If you don't want to be there, you have to commit fully to a distributed model.
For turbopuffer, distributed doesn't mean never meeting. Everyone gathers twice a year, recently in Banff and Mexico City. Its more distinctive practice is the "campfire": when a few people happen to be in the same place, they declare a campfire and invite anyone who wants to come. The week of the talk was a campfire in San Francisco, with customer meetings and dinners. Attendance is optional. Some people want to stay home, focus, and see their families, and only attend the two offsites, which Eskildsen said fits the model perfectly. Others fly roughly every two weeks. He told of an employee in Ottawa who saw colleagues dialing in from a New York meeting room during a campfire, got such FOMO that she took an Uber to the airport and flew to New York to join them.
The company also created "turbocredits." Extracurricular contributions like a conference talk or blog post earn one, and a turbocredit upgrades your next flight to business class, which Eskildsen said further encourages time together. Engineers who choose to spend two days on a conference expo floor talking to customers, which he called taxing, also earn one. He said the credits may take on a life of their own: someone has already suggested a central bank, interest rates, and a betting market for them.
Closing
Orosz closed by observing that although many AI companies use turbopuffer as infrastructure, the conversation had been mostly about engineering principles, curiosity, and how people come to trust each other and work together, rather than about AI.
I'm excited to have a chat with Simon Eskildsen, founder CEO of turbopuffer, a very technical CEO and we're going to have a pretty technical discussion. But before we jump into it, Simon, I wanted to ask where did you fall in love with computers?
Through PowerPoint.
PowerPoint.
I don't know if any of you know this, but in power well, you probably know this, but in PowerPoint, right? You can make the diagrams and stuff when you click them go to another slide. That becomes Turing complete real quick, right? You can sort of, you know, create very complicated convoluted games and then at some point you know, you make it through the Microsoft Office suite and you discover FrontPage. Do you remember FrontPage?
Yeah, I remember FrontPage. It was supposed to eliminate the need for all any front-end developers.
Exactly. And it only worked in Internet Explorer. I remember heartbreak I had one day when someone opened a website I created in Firefox and it was all over the place.
And then one day I accidentally clicked the HTML thing in FrontPage and it just showed all of this stuff that I couldn't make sense of and I just started looking at it and then going online and finding little snippets that you could add in to make the cursor change and all of these different things and then it just sort of escalated from there. Then you upgrade to Dreamweaver and now you're code and then you're like, well, how do you make the pages dynamically? You learn PHP and then for me I exhausted the internet on Danish language programming advice.
And I was around 11 or 12. And so I just, you know, went and got addicted to World of Warcraft for 4 years, but that gets you really really good at English.
So you kind of start hacking, get into deeper. Now the logical step would have been to just, you know, go to university and learn properly about this stuff, but that's not what you did, did you?
I mean, I just started just I mean, you know, then I learned video games, then I learned English, and then, you know, this like massive arsenal of the web. Now it'd be very interesting cuz the LLMs would just speak Danish to me, and you wouldn't have hit the wall like I did. So, that would have been very interesting. Maybe I would have been better at programming. That would have been nice.
And then I Yeah, then I just started picking up jobs and things like that throughout high school. And when I was in high school as well, I got exposed to this thing called the International Olympiad in Informatics. Heard of this thing?
Yeah.
And I had an internet friend, and she lived in Australia, and she was on the Australian team. And she told me, "Oh, there's probably something for the Danish team as well," but I'd never heard about it before. And so, I found it on some like little mysterious website, and then applied, and then solved these programming problems that looked very different from the HTML and PHP things that I'd solved until then.
What were these like the algorithm-ical-ish programs?
Exactly. This is not actually the kind of problem that you would see there, but I think it illustrates well the kind of problem that you might get, right? You can imagine something like, "Okay, here is like N trucks, here's M packages. The M packages have these dimensions. Give me which trucks, which packages should be in, right? And then do something optimal." Like that's an NP-complete problem. You can't solve that, but you could compete with everyone else in the competition of doing the best thing. So, it's these kinds of problems, right?
And so, I started doing that in high school. I was working as well for a startup. And then Shopify found me while I was still in high school.
And the whole like Shopify found me, was it through your open source contributions? Was it something else?
It was because I had written an article where I had dropped my iPhone and it was, you know, there used to be a time, right, where you drop your iPhone and you just knew it was over for the screen.
Yeah.
It doesn't really happen as much anymore. Like the screens have gotten a lot better. But back then it was like, yeah, one drop and it was dead and it just like couldn't use it anymore. And so I went back to one of these old Nokia brick phones. And this is back in 2013. And people hadn't really realized all the pernicious effects of smartphones at the time. And so I wrote this article about how, oh my god, I'm like calling people. And I have my sense of direction back. And I wrote an article about it. And this article it went on Hacker News briefly and New York Times decided to feature it.
No way.
Yeah. And so a lot of traffic was driven to it and then some astute Shopify recruiter put it all together. And I had a call with them and then I don't think they realized that I was still in high school, but I had a great call with them. They invited me on site to Ottawa, Canada. I had no idea what Ottawa, Canada is. I think the email says something like, "What's an Ottawa?" I had no idea.
And so I went there and it was just like walked into the building and it just felt right. And so I interviewed with them and then said, "Well, I got to finish high school first." And then I moved to Canada to work at Shopify. Yeah, in 2013.
Yeah, I think that's a like legit excuse for like not even worrying about college and university. But I did your mind?
It did. I thought I was doing a gap year. I thought I was like, "Okay, I'm going to go work at Shopify for a year and then I'll probably go back and do" But I was very insecure at the time about the fact that I hadn't studied computer science and my only exposure had been all the IOI competitions. It's a pretty good crash course in a lot of computer science. And if nothing else it had really taught me that you can just sit down and read a paper and just figure it out if you spend enough time on it.
So, I did that repeatedly and in my first year at Shopify, every time I heard something that I didn't know what was, I noted it down on a piece of paper and then I went home and then that evening I would just read about it because I felt insecure that like, well, if someone mentions like TCP, surely they know exactly what's in the three-way handshake and how TLS is like layered on top and they've looked at Wireshark and all of that. I don't think that's true, but that's what I thought. So, I went and did that for everything that I encountered. So, that was a really good crash course and then very quickly it became clear that well, I just want to continue doing this. I don't want to go somewhere else and then come back to this cuz I felt like I'd already found what I wanted to do.
So, it sounds like it was a pretty good combination of like you just having this like very natural insecurity. Like you're young, you know you don't have the education that everyone else has and inside the company they're just doing pretty like cutting edge stuff even at the time and even today, right? Like they're leading. And so, you just kept self-teaching yourself like just catching up and do you understand that you just went deep in every concept that you understood? You didn't like just try to understand at surface level, but like go as deep as you can, search on the internet, buy books, whatever that is.
I think it was just that I just wanted to keep learning how computers work and I think that this is something that I now look for when we interview engineers is that you can't help yourself but trying to peel back the layers. And for me, that ended up with the infrastructure layer. That was, you know, the people closest to the metal at Shopify. And I would just always sit next to them at lunch cuz I was working on the product side, but I just couldn't help myself. I just wanted to learn what it was when they were talking about a reverse proxy. I'm like, why is it reverse? I still can't answer that.
I mean, okay.
Do you know? Well, what's in reverse?
Because it's a proxy. Right?
I don't know. I don't know. It's like an inverted index. Like, what's inverted? It's a terrible name. Anyway.
Yeah. I mean, it's still better when you could do the NAT tables, the lookups, some of those things. But, yeah, I hear you. There's some like weird names with this. But, at Shopify what were some of the kind of like hard engineering challenges, outages, like learnings that kind of defined you that were really also fun at the time or interesting to learn, but it would have been hard to get it elsewhere?
Yeah, so I think it was, you know, in the 2010s there's like a bunch of SaaS companies that scaled really quickly. And I felt so fortunate to have a front row seat to that. And so, I ended up on the infrastructure team. It was back in, you know, '13, '14, and Docker was coming out. And so, we were containerizing everything. And every single year we had to you know, the growth rates of SaaS sometimes seems quaint in comparison to the growth rates of companies today, but it was a company that was growing at, you know, 120, 140% year over year.
And so, every year we were just preparing for a Black Friday that was going to be a lot worse than the last. And this is back in the day of we're buying physical hardware, right? We have to like place an order at a particular point in time and do some interpolation based on that. And the software also had to scale. And when you're scaling most software, a lot of the application layer problems end up back at the database layer.
And so, I just naturally found myself at this layer between Rails and the databases. Shopify didn't at the time at least contribute many patches to the databases themselves, but mostly just spent time orchestrating. So, we were doing sharding because as my dear boss Camilo used to say you can't cache writes. So there's a fundamental point where you just have to move beyond a single shard. So I joined around the time they did the sharding and I think they did the cutover a week before Black Friday which is mind-blowing, but it worked.
And then the subsequent years we worked on things like going into multiple data centers. We also had this big mysterious Redis server that was like you know 128 gigabytes of RAM which was a lot at the time, today it's not that much, and no one really knew what was in it and then it went down one day and people were like well that's super terrifying because people had just been treating it as this KV store and so we started splitting it out.
We did all this stuff around making sure that if you go visit a Shopify store and the thing that stores your sessions is down, the right behavior is not just for everything to be down but that's kind of the default failure mode, right? You're not going to rescue all of that unless you're in a programming language that really forces that decision. So we did things like build this matrix out of okay well this service when this component is down should act this way.
And I find myself writing the test suite for a bunch of that and then I was like okay well we can't just mock all of this and so I came up with this idea at the time of like oh what we're going to do is we're just going to shell out to GDB and then enter the process and then close the file descriptor to the database to simulate deep through the entire layer that the database fails. That was a little crazy and we never shipped that on CI but it did uncover a massive amount of issues in Rails to be upstreamed and things like that of just like handling failures at the connection layer.
So then I moved on to create this proxy called Toxiproxy and have you heard of this before? Yeah Toxiproxy is just like a layer 7 proxy that sits in between you and well, layer four, but in between you and the databases. So, you basically have just like this proxy and then MySQL, whatever, doesn't speak the protocol, but then you can do an API call say, "Take the database down. Make it slow." And over time it also added layer seven things of like do a bunch of failures. This way you're not mocking the low-level drivers, but you're testing the drivers and their failure handling as well. So, then this entire matrix could be implemented in CI.
So, basically the proxy was just like a really thin layer which like was passed through, but you built the functionality to like simulate problems with database or things like data corruption or whatever you want it to do, so you could just do it in there and then anything that built on top of it. Oh, yeah, and then everyone had to like call this proxy or it needs to be in a layer.
Exactly. So, you could do like MySQL, you know, Toxiproxy.MySQL.down and then pass it a lambda of what you want it to do, like get this page, do a checkout, whatever, with the sessions table down. And this just uncovered tens of issues, right? In the MySQL driver, in Rails, it's just like no one in the ecosystem had been testing for this and it was very difficult to see this in prod, right? Because when MySQL is down you're focused on just getting back up and not like what could the application actually have done.
Yeah, so it's interesting. Of course, we're going to talk a bit more about databases, obviously, but just thinking about how a lot of the problems or some of the most gnarly problems in large systems are always to do with state. And I never connected until now that state is usually there's a database. If there's no database, if you have stateless services, you know, I mean you still have problems, you have nodes going down, you have I don't know, corruption, whatever, but it's usually like more isolated. But basically like if we have state, we typically have databases. If we have databases and if you can simulate these problems suddenly you can like predict a lot of things. So, the problem with state oftentimes is it's really hard to simulate problems happening ahead of time unless when they happen. So, it sounds like you had pretty good success with
Yeah, I think to my knowledge it's still running in like the CI system of Shopify today. I don't know if anyone in the crowd is from Shopify, but I'm pretty sure that it still does. And so we wrote all these tests against it to implement all of these different failure conditions. And it just Yeah, it worked out great.
So, you spent 8 years in total at Shopify. So, like starting from like all right, just a gap year, I just went down a year, another year. At what point did you think about leaving and why? And what was your kind of decision framework? It sounds like you were like on an epic run and even today Shopify is doing wonderful. It's probably doing even way better than you're like, you know, that growth kind of kept on. So, I'm sure there would have been arguments to stay and you know stay on their flagship.
Yeah, so I spent 8 years there from '13 to '21. And I think there just came a point where I wanted to see something different. Again, I'd been inside of Shopify since I was 18 years old, right? I'd seen one other startup in high school. I was like if I want to learn more about computers and learn faster, it might be time to inject some novelty into this function. And so I left in '21 and I'd worked on so many different parts of the infrastructure like caching, me and Justine who's now my co-founder, we wrote the entire storefront for Shopify, which powered almost 100% of traffic 18 months after we embarked on it. We've worked on running Shopify in multiple data
centers. We've worked on so many database scaling projects like caching, all of these different things, right? A lot of the scalability came from the Kardashians launching lots of products on Shopify, which would force a lot of traffic. But that's eventually how I left.
And so when I left I didn't really know what I wanted to do, and so one of the projects I had while I was at Shopify was this napkin math project. Have you seen this? Napkin math? No.
So napkin math was essentially just this table that I maintain on GitHub of how much bandwidth can you drive to DRAM? What does a round trip to S3 cost and how long does it take? How much bandwidth can you drive to an NVMe SSD? How much bandwidth can you drive to an EBS volume? Just a collection of, there's probably like 50 of these numbers, and then a Rust script that generates them all.
What all these things cost? Like what does a gigabyte of memory cost? $2. What does a gigabyte of S3 cost? 2 cents. What does a gigabyte of disk cost? 10 cents, right? What does it cost on spot? What does it cost on a 3-year commit? I just have a massive table and then create flashcards for almost every single cell so I know all these numbers. And this was the project I started taking on at Shopify because I found myself in this role a lot where I would go in and review a project, right? So some product team would be like, okay, we got to build this thing, so we got to build this infrastructure to support the feature.
And a lot of the times they would say, okay, well, we've gone and benchmarked it on database A. But the benchmarks are not very good, so we're going to go with database B. And I hate benchmarks so much because that's not a satisfying answer to me.
To me it's like this does not jive with my intuition. Database A that you're saying takes 10 seconds to do this should take 10 milliseconds if you do the napkin math, right? If it's a search query, right? It's like, okay, you're searching for three terms, each term has this many documents that match it, that's this many megabytes. We intersect this many lists. You have DRAM bandwidth on multiple cores of 100 GB per second, this should take 10 ms. You tell me the benchmark takes 10 seconds. One of us is wrong. Either there's a gap in my understanding, which is very likely, or you benchmarked the wrong thing.
And in some ways, for some reasons, right? It's like, okay, you've done a benchmark, you didn't realize that your benchmark is doing a distributed query across 100 different nodes. And so of course the P99 is going to be really, really high, right? Unless you've cut that off or made some different set of trade-offs. So I just found myself in these discussions repeatedly where people were making infrastructure decisions based on poor benchmarks.
And so I needed some ammo to go in and just be like, okay, we can just do the calculation right here, because I was always doing these little demos or writing little prototype scripts to demonstrate this. But it was just the argument of: here's how a B-tree works, this is how many pages we have to visit, this is what a random SSD read takes, it takes 1 ms, you have to visit 1,000, blah blah blah, and then present it back and say, this is the difference to your query. Well, is the query plan correct? Is there a bug in MySQL? Do we have bad disks? What's the discrepancy here?
And I just got caught with that bug. And so after I left Shopify, I was just writing a lot of articles about this. I was just like, well, how long does your disk query take? And then one hypothesis that I had at some point is like, okay, well, how many writes per second can MySQL do? Well, shouldn't the amount of writes per second that MySQL does equal the amount of fsyncs that you can do per second? That sort of makes sense, right? Every time you do a write, you fsync to persist to disk. So, how many fsyncs can you do per second? Well, an fsync takes 1 ms, so you do 1,000 writes per second. Well, that doesn't really match up. I feel like a database can do more than 1,000 writes per second. Why can it do that? So, that was one of those things where I tested and it's like, okay, well, MySQL on a little dinky box could do 10,000 writes per second. Well, how is that possible?
Mhm.
And now you would just ask—
How is it possible?
Because you batch. So when fsync happens on usually a 4K—
Yeah.
Right? But it's like that's not intuitive. I got caught, it was probably some 24-hour period where I just got obsessed with this question, where you're writing the BPF traces and all of that to do all of this. This is pre-LLM, so it took forever. And then I found out that, oh, every fsync was much larger than I would have inferred. It's like, oh, it's batching. You go into the code and you read it, and then you find some obscure article by, it's always someone in like a central German town that's written some article about how some intricacy of MySQL works and a patch that they did to it. The entire internet runs on small towns in Bavaria. I'm convinced. Yeah.
[gasps and laughter]
And then you decided to start turbopuffer.
Yeah.
How did you decide? Did you know what you wanted to build, or was it more like I want to build something something databases, cuz you were clearly very into databases. You've done an awesome job benchmarking, like what are the theoretical limits? You are very familiar with this, probably became, you know, a world expert in this niche. And then how— Did you want to go into databases again?
I think it was— There's three things that sort of came to a head. The last project that I worked on at Shopify was search. And I didn't have a good time.
What did you use back there?
We don't need to name names of other database companies, but it was one of the traditional search companies that a lot of different companies run. And it was just very difficult to get it to do what I wanted, and the projects that touched that database, I couldn't get them to perform at the napkin math. And there's no query planner, and I couldn't figure out why it wasn't there. And sometimes it tracked, and then sometimes it really didn't track at all. And so I tried to learn as much as I could to figure it out and started reading the source code of it, and I just couldn't get it to track very often. It was very difficult to operate. And so that was sort of in the back of my head. I never thought I would touch that again.
Then the second ingredient was the napkin math project. Because it sort of just gave me a lot of facility with all of these napkin math numbers of what might be achievable with the machine if you utilize it perfectly properly. Yeah.
And then the third one was that, leaving Shopify in '21, having spent eight years there. And during that time I did this thing I called angel engineering. So I joined my friends' companies and then I just vested equity instead of just investing or something like that. Cuz I wanted to have my fingers in it. I wanted to see what else was out there. That's why I left. And this problem kept coming up again and again, right? ChatGPT came out in 2022, and I was working with a company then. And they wanted to connect a bunch of documents to AI, and that's when the context windows were really small. So you had to reach for search very quickly.
So it's like a few kilobytes.
It was 8 kilobytes or 4 kilobytes depending on the model. It was very very small. So you had to reach for search very quickly, right? And so I worked with them, and I created a little recommendation engine. And the recommendation engine was actually quite good. I found out that one of the co-founder's wife was pregnant through the recommendations that I was getting when I was running it on his feed. Like it was—
Weird, but yeah.
recommending. Yeah, I mean it was just like, you know, he was reading about like— I did get permission. I just don't think anyone expected it to be good enough. And it's just like, okay, this thing is working.
And then I ran the back of the envelope math on what it would cost to do this for everyone, like all the users. This is a company called Readwise. So there's articles that you save and then search later. And it was going to cost 30 grand a month. And this was a bootstrapped Canadian company. At the time they spent about 5K a month on all the other infrastructure combined. So it just didn't— fundamentally in a company, if you're doing an investment, you have to run some gross margin on top of whatever you're paying, right? And it just didn't line up.
And so we just didn't ship it, and I worked on, you know, tuning autovacuum on Postgres or something like that, which is a good pastime. And then I just couldn't stop thinking about why it was so expensive to store all of these vectors that we were using for the recommendations.
And I just sat and did the napkin math one day of, can we just put it all in S3 and do some clustering and then organize the files in just the right way, and it's like, maybe you could build that. And then one day I just kind of said [ __ ] it and did it, and sat down and started to write it out. And I spent the summer of '23 just hammering my head against the wall trying to find an approach where I could get the latency that I wanted. Because the problem with S3 is it has really good durability, but latency, we're talking hundreds of milliseconds, right? Yes, the P99 on a 256 or 512 kilobyte object on S3 is around 200 milliseconds.
And then you're saying P99 cuz when you're talking large scale you want to care about the P99, right? Yeah. That's why we're not talking about P50. When you're designing a system you want to optimize for the P99, and especially because when you're designing a system on S3, generally in every round trip you're not doing one request. You're often doing lots of requests, right? You're going to hit the P99 real quick. Exactly. So it's like if you're navigating a tree on S3, right? Okay, you get the upper layer of the tree, 200 milliseconds, you get another layer of the tree, 200 milliseconds. You get a bunch of leaves of the tree, it's 200 milliseconds. So in aggregate you want to look at the P99, probably even the P999, to design the system properly, cuz you would need to minimize the number of round trips that you had to make.
So, I just sat and sketched that out and tried a bunch of different approaches, and then finally in July of '23 I got something end-to-end that seemed to work, and then rewrote it probably twice, and then released it in October of '23, based on just that summer of working through it.
And then you built it on top of S3, cuz I guess durability and all of it is just really good. How did you make it fast?
We didn't in the beginning, or I didn't in the beginning. It was just me at the time, and it was a project. It was not a company. It was to satisfy a curiosity. I did not set out to do this as, like, I'm going to go raise $10 million and do it. I barely knew what a VC was. I just had to do this thing, and I was so focused on doing it. It was so clear to me that if I wasn't going to do it, someone else was going to do it, and I just became fully obsessed that summer with it. And so the first version was the simplest possible thing.
I think I'm a very pragmatic person. I barely read any of the literature on LSM. I sort of, you know, read a bunch of it and got the basic idea. Barely implemented that, because that would have taken too much time. It was the simplest possible version of what it could be. Really what you have to imagine is that the simplest way you could do this is you run some clustering algorithm on the vectors. You get the clusters and then you put the clusters in files. The files are called cluster one, cluster two, cluster three. And then you have another file called centroids of the clusters, and then you do the search by downloading centroids, looking at the centroids, and then downloading the N closest clusters. There's a few optimizations around merging some clusters that were adjacent in files and so on, just to control some costs and some performance, but that was basically it, and then getting that to scale. That was the first version.
And then how do we make it fast? Well, I didn't even implement a caching layer. I just put a reverse proxy in front of S3 with nginx and then had it cache it.
Do you know what a reverse proxy is?
I can— I do know what it is. I just still don't know what the reverse is about. But anyway, the reverse proxy reverses things. The performance in this case, maybe that's what it's about, by caching, right, all of the S3 objects. Again, it was as simple as, I'm just going to put that in front. I knew how to configure nginx. I've written more nginx Lua than— a lot of nginx Lua, very good software. Just had a cache in front. And then the way that I would do things like deleting in the cache was just shell out to xargs and remove things in the cache and reverse engineer the directory structure on nginx. And that's what we shipped, and it was just running on a single server in a tmux instance. I was like, "Okay, let's see if anyone gives a shit."
Yes, so so far, I mean, this is kind of cool engineering and a cool side project and a bunch of novel ideas, and I think just some hardcore engineering. How did Cursor come into play? Cuz when I learned about turbopuffer, I was talking with Cursor about how they built their back end, their database, how they scaled, and they're telling me all these migrations, and they're telling me like, "Oh yeah, so we were on Postgres, but it didn't—" No, they did something else in Postgres. It didn't really work that well. They went to AWS Aurora, which is AWS's managed service for Postgres, and it didn't work well, which is very surprising. And they're like, "Oh yeah, and then we went to this thing called turbopuffer and it worked well." And I was like, "What's turbopuffer?" And they're like, "Oh yeah, turbopuffer," I think they said we were one of their first customers. And this never computed to me. Cursor was already massive at that point. How did you meet the folks, and how did they become— Were they the first customer, one of the first?
They were the first customer.
The first?
The first.
No.
They reached out after I just launched on Twitter. I was like, "Hey, I built this thing." And frankly, it was like, "Hey, launched this thing." And to me it was like, "I am so sick of working on this." I was like, "I've been working on this all summer. I don't know if anyone cares. I only want to work on this if anyone cares. Let's put it on Twitter." Again, a single tmux instance on an eight-core node somewhere in GCP. I was like, "If someone goes to prod, I'll set it up properly on multiple, and I'll just block on that, but let's see if anyone cares." It was like the MVP of MVP. Anyone who's actually worked on the internals of databases would have had too much pride to ship anything like that. And I was just releasing it like a SaaS project. Why can't you work on a database like it's SaaS? Do you know? It's like, if anyone uses it, we'll do it properly. I know how to run software with a lot of nines. But it was not a proper L set. Like
It was the simplest version of what it could be. And then I released it on Twitter. I was like, "Yeah, you can do a million vectors for a dollar." And before that, I think the cheapest was maybe $100 per million for something that actually worked.
Yeah.
And I knew it was reliable, right? I knew like I had to use invariants, like if you shut down all the VMs, like no data is lost, like all the writes are committed directly to... like it has all the same invariants it had today.
And Cursor reached out. And knowing them now, I'm sure at the time Cursor was maybe eight people, and knowing the founders now, I am sure that they had sat at the dinner table one day and were like, "The unit economics of what we have right now, where all the vectors are in DRAM, are not working. Why hasn't anyone built it where we can put it in S3, and the actual codebases that are actively being used, we can put in memory?
And everything else just sits in object storage, and then we just hot load it in and out of the cache." Makes so much sense, right? You open the codebase, a few seconds and it's in RAM, and then the queries are as fast as anything else. It made so much sense. So, I mean, at the time they were... if you look at some of Aman, one of the co-founders', early tweets, he talks about using S3 for KV caching and things like that, which barely anyone is still doing even though the economics are economical, yeah, price-wise.
Yeah, and it's very uncommon, and I think it will happen, right? But they were ahead of their time. And I think they were... I don't know if
they were thinking of building it themselves. I think that's quite likely.
And they found turbopuffer, and it just perfectly pattern matched into that. Again, I don't know if this dinner conversation happened or if this was just inside Aman's head. But it pattern matched something. And so, we exchanged a bunch of emails, and then something compelled... I didn't know anything about B2B sales. Now, I love B2B sales.
I didn't know anything. I was just like, I just want to help them, because they had some unit economics that didn't line up. So, I just went to San Francisco, right? I live in Canada. I went to San Francisco, and I showed up at the office. And when I showed up at the office, they were having some Postgres problem that they were discussing.
Yeah, they had the Amazon Aurora problems, yes.
Yeah, early on, and I was like, "Oh, do you guys have pganalyze?" And they said, "Oh, no, we don't." I was like, "Okay, let's get that going, right? Let's look at it." And it was the same thing as it always is with Postgres, which is autovacuum hadn't run enough, and so they had all of these, like, going to heap when they should be doing index scans and blah blah blah. So, we were talking about all of that. I was just helping them, right? It was like my, you know, my database genes just kicked in. And I think this built enough trust with them that, okay, well, maybe if he knows how to help us with the database, maybe he also would know how to build one.
And at this time I'd also approached what I thought was the best engineer who ever worked at Shopify, my co-founder Justine, and she'd come on. And the first thing that she did was replace the reverse proxy NGINX cache with a file-based cache. Just a direct cache, which, again, great. Like the S3 thing worked. And so she was online. She was starting to work on it, and Cursor then that night was like, "Okay, well, we're going to migrate." And so they migrated everything over the course of like a week or two after that.
But Cursor was a small company back then, right?
Yeah, and they were just in the beginning of their massive rapid growth.
Exactly. And I told them that I was going to reduce their bill by 95%. And I did. Like, we did. Justine and I did. They came on, and their last bill with their previous vendor and the first bill with us, it was 95% lower.
Yeah, and you're nice for not saying vendors, but I can say vendors because I talked to them, and it's in the deep dive about Cursor. It was AWS Aurora specifically. So
This was not Postgres, no. This was a
It was a different one, but it's probably still in the write-up. We don't need to name names, but yeah, the reason they won there is reliability was their main pain point. I'm sure the unit economics would have been there.
But yeah, this was... And then what Sualeh told me is he said, like, "Look, there's a few things that we did that you should never ever do." And he said, "One of them: you should never ever bet your business on a tiny startup where you are their only or biggest customer, except for turbopuffer." And he said, "I love those guys." So I guess it just comes to show that even in your case, like, to me what this story shows is you can do things when you build high-quality things and you're pushing for things, good things can happen. And on the other side, for Cursor, when you're a startup, it's okay to take sometimes irrational risks when you have conviction. And it sounds to me that you gave them conviction by showing up in person, by helping them, by showing that, you know, you know your stuff. Like you suddenly brought in your 10-ish or 8 years of Shopify experience and your curiosity. And they probably took a risk because of that. Not because you were some, you know, random vendor. They probably would never have done that.
So, fast forward to today, turbopuffer is now a lot bigger. You're working on some cool things, but you have this very interesting business where for you CPUs are important, right? You run on mostly CPUs. And you told me a story over dinner yesterday that you met Jensen, and Jensen really wanted to sell you on GPUs. Can you tell me how that meeting went?
Yeah, Jensen Huang, right?
Yeah.
I never met Jensen before. We were at an event at Nvidia, and we were just doing presentations.
This is in their big HQ, super impressive.
Yeah, exactly. They invited a couple companies to go and talk about our businesses and how we can partner with Nvidia and so on. And I don't know, I think I was in a goofy mood that day, and so I went up on stage, and I said, "Hey, I'm Simon from turbopuffer, and yeah, if you're wondering about the name, it's like if everything goes south, we can always pivot into vapes."
I was kind of nervous. This is what I said. And then he said back to me
Wait, who was in the room? Was it Jensen? Was it a direct report?
It was Jensen, and then I don't know if it's just 50 direct reports, or it was like, you know, it was Jensen, and then a bunch of the Nvidia leadership, right? Because you go there, and then you talk about... you find opportunities just to partner and work together, right? And so I said, yeah, you know, so plan B could be that we could pivot into vapes. And then he said... I was already nervous. He said, "Judging by your slide, maybe you should."
No, he did not.
I didn't know what to say back to that. So, I said, "Well, Jensen, do you vape?"
He didn't answer the question.
Someone on the team wrote to the whole company, the turbopuffer company, "Simon just asked Jensen if he vapes."
And then, you know, this is a great start, right? And then the team had sort of talked to me beforehand. It was like, "Simon, we've got to make sure we don't say the C-word. We can't say CPUs." So I just couldn't stop talking about CPUs. I was like, "AVX-512 is so sick. Like, we love SIMD, and like, there's so many CPUs. They're so easy to get. Like, it's just a riot in CPU land." Like, you know, I think I stopped short of saying, "I'm so glad I don't need GPUs," but I just couldn't stop talking about CPUs. Yeah. And so, you know, Jensen took an interest in that. Yeah. So, who knows? Like, I'm sure he may remember a little person, maybe
He made it his mission now to, like, at some point get you guys onto GPUs. But speaking of CPUs, can you tell me a bit what you're seeing inside of the hyperscalers, the cloud providers? You're now on AWS, you're in GCP, you're on Azure. What I would think naively is there's a GPU shortage, and when I talk with inference companies and AI labs, they're just getting whatever they can. I would think getting CPUs should be easy. Is it?
No.
It's not anymore.
Why? Well, what's happening? Can you tell us about the dynamics on the why and what you've learned?
Yeah, so I think that GPUs will probably continue to be scarce. Like, I don't know, maybe there's going to be some surplus. I refuse to speculate too much about the macro, but I think as RL is becoming a very, very large amount of the workloads, that needs a lot of CPUs. So, the labs are sucking up a lot of CPUs, because you need CPUs to be like, okay, we need to teach this model how to search. We need to teach it how to use grep. We need to teach it how to boot up bash. It needs to run real things and learn from that. It takes a lot of CPU.
Mhm.
And so I think RL is consuming a lot of CPU. And then also just all of the agents are running on CPUs, right? They need to do all kinds of very general-purpose things on a CPU. And so as the demand curve is sort of shifting to the right and it's becoming more and more applied, and that feeds back into RL, by the way, right? Because as things become more applied, like, oh, the models are not that good at CAD or shipbuilding, I don't know. And then, you know, you have to spin up even more RL environments to do that. So, I think that's what we're seeing. And so we're on the other end of that, needing these CPUs. We need a lot of NVMe SSDs as well. And a lot of this right now is tied up in DRAM, right? Where you need a lot of that also for the GPU servers. But I would assume that it gets a lot worse before it gets a lot better on the CPU side. And I think even the big companies are fighting amongst each other, right? To get the allocations. And even we, you know, we're selling to companies that we also fight for CPU with and against, right? It's really difficult. And so you write things to try to make sure you get these CPUs as fast as possible.
Yeah, and yesterday I was at a dinner that you hosted with your team where you actually had a bunch of turbopuffer customers. A bunch of them are AI labs or AI startups, and one of them, Reflection, has a huge, massive footprint. And they were telling me that they're in a situation where they cannot buy more. When it comes to GPUs or CPUs, they max out. They have the longest contracts possible, and I didn't realize how competitive it is in the cloud when you go beyond a small fish to like a medium-size or even a large fish. It's interesting. So now you have this, and even you're having this kind of fight behind the scenes that is maybe not as visible.
Exactly. And I mean, you work with the clouds, right? You work with them to talk about which regions have CPU, which regions are getting... It comes down to power, right? Of like, okay, well, where is the power, which is generally where they're going to ship the new CPUs. And so we have to work with some of our biggest customers on that. So these are real constraints, right? That are making their way to us.
We're just very fortunate that it's very easy for us to run lots of turbopuffer clusters, because all we need are like a few CPUs and NVMe SSDs and then S3, and then we're in a good place. But there's lots of changes that we can make even to the architecture to try to protect from a lot of this. Now, I'd rather spend that engineering effort on other things, but we are very, very good at using a lot of very different SKUs, right? So we don't need everything to be a particular CPU or instance type. We can run with many different types of machine types on different
Meaning that's a fancy name for, like, the different machine types.
Yes, exactly, right? Like, you know, C4D or IAG or whatever they're called.
What's your favorite one?
We really like right now the C4s on GCP.
GCP, yeah.
The Z4Ds are also performing really well, now that we've done a bunch of optimizations to them. Those are really, really great machine types. We really like those. And then the Arm C4As as well on GCP. We like those, but I think that in general, when you're small, it's very easy to suck up a bunch of... But at Shopify, I was also part of, you know, deciding ahead of BFCM, right? A few months out, you have to tell the cloud providers how much you're intending to use, do commits and all of that, right? The clouds are not as infinite as they seem when you're small.
Now, one way, of course, to get infrastructure and also just credibility is venture capital. If you raise $100 million, a billion dollars... some of your customers just raised $2 billion. Actually, I talked with them yesterday. You know, it gives you credibility, it gives you cash, you can pay for this thing. Your specific, turbopuffer's relationship to venture capital seems very interesting. I never heard you announce a raise until maybe just very recently. And you told me that when you started this thing, you didn't think too much outside of just building some cool stuff. How did you think about venture capital? And how do you think about raising? Because again, I feel you have a very fresh and different perspective than what is typical inside of Silicon Valley.
Yeah, so I think to understand how I think about capital, you have to go back to the beginning of turbopuffer, right? Where I promised Cursor that Justine and I could get their bill to 4K a month. And this was based on some very rough napkin math on, "Okay, if turbopuffer was a better implementation than it currently is, then it should cost this much." And that's the price that we shipped with, and that's what we guaranteed Cursor. But the software was not that good. Like, it was very reliable, but it was very simple, right? And that's a core engineering principle for me: simplicity above everything. You and I have talked before about software that ages well and some of the advantages of having long tenures inside of companies. You had a long tenure at Uber, I had a long tenure at Shopify, so you see simplicity just almost always wins.
And at the time, I was not convinced whether this was a venture-scale opportunity. Because I understood that if you take venture capital, no matter how many smiles there are in the room, everyone's sort of expecting that you have to earn a big return on that on some timeline that makes sense to everyone involved. And everyone involved are, you know, pension funds in Canada. Like, it is a whole stack, right? Of people that need to... So, at the time I was like, you know, I don't know if this could be a billion-dollar company. I didn't know that in the very, very beginning. It wasn't completely clear to me. It felt like a very niche kind of product, right? To build this particular search engine. And that was completely fine with me. So, you know, it was fine. And so then I just looked at the Cursor bill and I looked at my GCP bill, which is what we started on. And, you know, I was like a dumb Danish person who's just like, okay, this number should just be lower than the other number.
Yeah.
That's sort of like, you know, and it's just I don't think I'd spend enough time in San Francisco cuz I think the money over here, it works a little bit differently. That's just, that's all I knew.
You were doing business 101: as long as you're making a profit, you're good, right?
Yeah. Like that's, it's like I'm not kidding in this exaggeration that it was just like, that just made sense to me. That Justine and I were just going to go optimize this until these numbers were roughly equal. And maybe if we could get some other workloads, we could start paying ourselves. But that was like very much the philosophy at the time, because I didn't know if I could go raise a bunch of money. I didn't know anyone who had the money. I didn't have any relationships.
You're an absolute outsider to the...
I was an outsider. I was like an outsider squared, right? I grew up in Aarhus, Denmark and I then moved to Ottawa, Canada. So, it's like I'm an outsider to Canada and in Canada I'm an outsider to San Francisco. So, I was just thinking about this from first principles. Like, oh, you're a venture capital, you need this return, you need it on this timeline. I don't know if I can deliver that yet. I would need more data to decide that because I wanted, like I kind of want to keep working on this and now I have to get to this point for it to not be a failure.
In January then, there's a person that I was at IOI with in 2012 and 2013 and his name is Boyan and he was on the North Macedonian team at IOI. And he was really good. He was so good that the North Macedonian team called him God. I don't know why, but that was what he went by and he was, yeah, he was very good. And I really wanted to work with Boyan. But I couldn't afford to work with Boyan.
And he was very much like, this is what I can live off. Like, you know, I just want to build this, like that would be, like this is what I can do. But at this point Justin and I hadn't taken a salary for like 6 months and we'd already spent like tens of thousands of dollars on GCP bills and all of that and I was like, I don't think we can do it.
And so, I had met one individual in Silicon Valley, his name is Locky, and I ended up just calling him and saying, hey, I kind of want to learn a little bit faster here. Can we raise like 700k? That's like what I wanted to raise. So, this is like, I want to have two engineers for the rest of the year, Justin and I still don't need to be paid, and then a little bit of buffer room. This is what I need. And if this doesn't have PMF and is a big opportunity by the end of the year, I don't think we're going to bother and we'll just shut the whole thing down and we won't have to take in a dime, we'll return everything to you.
I think that was the first time he'd heard anyone say it like that. And I told some other VCs that at the time and that was terrifying to them. I think to someone on the West Coast, this sounds like you have low ambition or something like that.
Mhm.
And to me it was just like, I don't know, it just came from a... When I don't know how to play a game, I just play with open cards. Like this is how I see it. And so it was very clear to us that we wanted to do this, but also it became clear to us that we didn't want to just keep working on this unless it could become big. And we were starting to develop conviction that this could actually become really, really big.
And so we did that and hired Boyan and then became profitable later that year. And then just continued to hire and then it's like, to raise more money, you need sort of... There's six reasons to raise capital. The first reason to raise capital is to fund R&D.
Mhm.
That was the reason that we raised capital in January, because we funded R&D with a lot of our own, you know, opportunity cost in not taking a salary and then paying the bills ourselves. But we wanted to learn a little bit faster and so we hired Boyan and Morgan as the first engineers.
The second reason to raise capital is to fund growth. You've built something and you want to tell the world about it and you want to spend more capital to do that.
The third reason to raise capital is for the founders' ego. It's a very popular...
The honesty.
It's very popular. Very, very popular, right? Big numbers, lots of press. And I think this is a very, very dangerous reason to raise money and I wish that it was more talked about, because you're diluting all of your employees when you do it. You are setting a certain price for future employees and their upside. For some people it can become a status game and that's not what it's about. We're here to build a big business together and this is not a reason to raise money, but I do think that it happens.
The fourth reason to raise capital is to reward your employees, right? You're on a very long journey and you want to work with the best people in the world and by definition there's not that many best people in the world, so you want to reward them. That was the reason that we took more capital in December, was to allow the employees to liquidate some of their equity instead of waiting for some event, like an IPO or something, further out.
The fifth reason to raise is for a strategic partnership. There are strategic partnerships that have been made in this city that have made companies. The sixth reason to raise would be doing M&A or something like that. But you have to be very honest about what reason you are raising for out of those six. The first time we raised it was for reason one and the second time we raised it was for reason four.
So which one was the first reason to raise?
R&D.
R&D, and the second reason was?
To provide liquidity to the employees. Yep.
I think it's a nice and healthy way and I think, yeah, the ego part we don't talk about, and the identity. Especially the closer you are to tech ecosystems where a lot of people are raising, it will be part of it.
As a closing, I want to ask you about the way you have a remote culture. These days I'm seeing, especially for companies that do anything with AI, may that be building AI infra or just AI products, a lot of them prefer in-person, having an HQ, oftentimes in SF or wherever the headquarters is, may that be London or somewhere else, because these companies often find that they have faster iteration. There are just fewer layers in between, and of course speed is very, very important. You started fully remote and you're still fully remote. How is it working, and what kind of quirks or turbopuffer ways have you found to make this work better?
Yeah, I think the company started in '23. So we're sort of on the cusp of COVID, where a lot of companies were just remote. The Shopify infra team was remote since the very beginning, cuz it's very difficult to get them all to move to Ottawa. And so it was natural to me, like, "Okay, I think there's kind of maybe two cities where you can build a database company fast." And that's San Francisco and maybe New York. There are maybe other cities, right? But that's kind of where it's been done. And so if you don't want to do that, I think you have to go all in on some distributed model.
And so we've tried to figure out what does that distributed model mean for turbopuffer? It doesn't mean the absence of in-person. We get everyone together twice a year in some location. Earlier this year we were in Banff, right? And then we were in Mexico City and so on. So that's not that uncommon.
But one of the things that we've been trying to do is we have this concept called campfires. And the concept of the campfire is that when a couple of people just sort of randomly congregate in a place, you call it a campfire and you encourage as many people as want to come to join. So for example, this week is a turbopuffer campfire in San Francisco cuz I'm here for this conference and a bunch of other things. And so everyone is invited to come. We're going to go meet customers, right? We're going to put on dinners for our customers and things like that. And we just make a thing out of it and spend time together. And we encourage everyone to come.
We've also gone to the extent now of... we want to encourage that, but not everyone needs to go to the campfire all the time. Some people just want to, you know, lock in and hack. And that's great. We have people that just make it to the offsites twice a year. And otherwise, they're home, they're with their families, and they don't spend time on an airplane. Fantastic. That is completely compatible with this model. And there are other people at the company who are on a plane probably every two weeks.
We had someone the other day where they saw a campfire happening in New York and everyone was dialing in from a meeting room in New York, and she had so much FOMO that she took an Uber straight to the airport in Ottawa and flew to New York to hang out with the team, right? And I think that's fantastic.
And we've also introduced these things where if you do a conference talk or a blog post or something like that at turbopuffer, something a bit extracurricular, we give you a turbo credit. And a turbo credit allows you to upgrade your next flight to business class. Which again encourages spending time together with the team.
And now, I mean, turbo credits are probably going to take on a life of their own. Someone was talking about doing a central bank and doing interest rates on the turbo credits, and doing a betting market on the turbo credits. And so this might take on a life of its own. And, you know, if you're at a conference like this, there are some of our engineers here who just want to interact with customers, and standing on an expo floor all day is quite taxing. And so if you do that for two days cuz you want to do it, oh, you get a turbo credit, right? And so it's just these fun little things that we try to do to encourage people to meet if they want to meet.
Thank you. Well, in this session, what I found very interesting is that so many AI companies are using turbopuffer as an infrastructure layer, but in this conversation we managed to talk very little about AI and a lot more about engineering principles, pushing, being curious, and the human connection, how important it is for people to work together, to trust each other. So just thank you very much for that. So let's give a big round of applause for Simon. Thank you so much.
This is great. Thank you.
Article published
