Inside Amazon S3: Consistency, Correctness, and Durability at Hundreds of Exabytes

Open on YouTube ↗
Overview

Amazon S3 looks simple from the outside: you put data in and get it back out. In this episode of The Pragmatic Engineer, host Gergely talks with Mai-Lan Tomsen Bukovec, VP of Data and Analytics at AWS, who has worked on S3 since 2013. The conversation asks how a system this large stays reliable and keeps evolving. Topics include the move from eventual to strong consistency, the use of formal methods to prove correctness, how failure is designed for, and why S3 now stores tables and vectors as well as objects. Tomsen Bukovec's recurring position is that S3's scale is only manageable through constraints, a design that assumes failure, and a strong commitment to simplicity.

27 min read

The scale of S3

Tomsen Bukovec gave the scale in numbers. S3 holds over 500 trillion objects and hundreds of exabytes of data. It serves hundreds of millions of transactions per second worldwide and processes over a quadrillion requests a year. The physical system underneath consists of tens of millions of hard drives across millions of servers, in 120 availability zones across 38 regions. The system is layered: disks sit in servers, servers in racks, and racks in buildings. The illustration the team likes to use is that if all of S3's drives were stacked on top of each other, the stack would reach the International Space Station and nearly come back.

The host said they had to look up the term "exabyte," which is a thousand petabytes, because a company with a few petabytes is already considered large. Tomsen Bukovec replied that some individual customers store exabytes in what they call a data lake. They mentioned that the Sony Group CEO had recently described Sony's data as a "data ocean" rather than a lake, and said that at exabyte scale the ocean is fundamentally S3. Their point was that most customers never think about this scale. They assume the drives are always there and experience S3 as something that "just works" for any kind of data.

Origins: unstructured storage and eventual consistency

The host mentioned a story they had read about a distinguished engineer in a Seattle pub who was frustrated that Amazon teams kept rebuilding the same infrastructure. Tomsen Bukovec did not confirm or deny the story. Instead they described the problem S3 was built for. Development began in 2005, and S3 launched in 2006 as the first AWS service. Amazon's engineers were building things like e-commerce sites and had large amounts of unstructured data such as PDFs, images, and backups. They wanted to store it at a price low enough that they would not have to think about storage growth.

The original design was built around eventual consistency. S3 would not acknowledge a put until it actually had the data. However, a subsequent list operation might not show the new object yet. For e-commerce this was acceptable: if an image did not appear immediately after upload, a person would simply refresh the page.

Tomsen Bukovec noted that Apache Hadoop also started as a community in 2006. "Frontier data customers" such as Netflix and Pinterest combined Hadoop with S3's properties, which Tomsen Bukovec summarized as unlimited storage, good performance, and a good price point. They extended unstructured storage to tabular data and built the first data lakes around 2013–2015. Tomsen Bukovec described these companies as born in the cloud. From about 2015 to 2020, enterprises adopted the same pattern. Around 2020, Tomsen Bukovec began seeing "a ton of exabytes" of Parquet files as customers applied S3's characteristics to tables.

Parquet, Iceberg, and S3 Tables

Around 2019–2020, Apache Iceberg began to rise. Iceberg gives table semantics to underlying Parquet data, and many of S3's largest data lakes adopted it across industries. Tomsen Bukovec explained why customers care so much about it. They want a "decentralized analytics architecture," in which different business lines or teams choose their own analytics engines as long as those engines are Iceberg-compliant. With Iceberg as the common format for tabular data, chief data officers and CTOs get future-proofing: they can replace analytics engines or adopt new analytics and AI tools while Iceberg on S3 remains at the bottom of the stack.

AWS responded by launching S3 Tables in December 2024. Tomsen Bukovec said the team added over 15 features to it in the following year. S3 Vectors entered preview in July and became generally available the week before the recording. Tomsen Bukovec described this history as a story that customers had written for data, with the S3 team following along.

The developer model: put, get, and new primitives

Tomsen Bukovec said the goal since 2006 has been a very simple developer experience. When engineers discuss what to build next, they return to the question of how to keep S3 simple to use. At its core, S3 is about the put and the get, and doing those well at scale.

The host added the other basic operations (delete, list, copy) and the core concepts of buckets, objects, and keys. Tomsen Bukovec explained that new capabilities have been layered on top based on what developers are trying to do. Conditional operations are one example. S3 previously added put-if-absent and put-if-match, and more recently added copy-if-absent and delete-if-match. These let applications make writes depend on their own logic.

"Object" is also no longer the only native unit. The two newest primitives are Iceberg tables (S3 Tables) and vectors. Under an S3 table is a set of Parquet files that S3 manages for the customer. Vectors are different. A vector is basically a long string of numbers, and Tomsen Bukovec called it a new data structure for S3 that sits in S3 alongside objects.

Pricing: removing the cost of keeping data

The host recalled that S3 launched at 15 cents per gigabyte per month. The host described this as a third to a fifth of the going rate at the time, which they put at around 50 to 75 cents. They had read that 12,000 developers signed up on the first day. The host also noted that AWS kept cutting prices, to about 2 to 2.3 cents today, and asked why the team did this when customers would probably have paid more.

Tomsen Bukovec answered by pointing to S3's mission, which they described as providing the best storage service on the planet. They cited IDC's estimate that data grows about 27% per year. The host remarked that this sounded low, and Tomsen Bukovec agreed that it is an average: many S3 customers grow two or three times faster, driven by sensors, applications, AI, and ever higher-resolution phone cameras. To keep all that data, customers need to grow storage economically. Tomsen Bukovec said S3 customers don't have conversations about which data to delete because they are running out of space.

According to Tomsen Bukovec, the approach is about total cost of ownership, not only headline prices. AWS sometimes lowers the price of storage and sometimes the price of capabilities. For example, it significantly reduced the cost of compaction for S3 Tables within a year of launch. It also offers tiering and archiving. Intelligent-Tiering, which the host said launched in 2018, watches access patterns. If data goes untouched for a month, it applies an automatic discount of up to 40%, so customers don't have to manage tiers themselves. Tomsen Bukovec tied this to how data is used: pre-training and fine-tuning models, analytics, and future uses customers have not yet imagined.

Glacier and designing to constraints across the whole stack

The host asked about Amazon Glacier, which they said launched in 2012 at one cent per gigabyte per month when the going rate was around 15 cents, in exchange for retrieval that could take hours. How was that trade-off possible?

Tomsen Bukovec described engineering as working within constraints, and said the constraints on availability and cost push the team to be creative. Because S3 is built "all the way down to the metal," including drives and hardware, the team can find efficiencies at every layer. Engineers set a target such as the cost of a byte and pursue it throughout the process. That process includes the data centers themselves and how technicians operate S3 physically, in the same way the team optimizes the software layers. In Tomsen Bukovec's account, this whole-stack view down to the buildings, together with close attention to the cost and lifetime of every byte, is what makes products like Glacier possible.

Why eventual consistency favored availability

The host returned to the original eventual consistency model and asked what it had bought. Tomsen Bukovec said the main optimization was availability, not durability. Consistency here means that a get reflects the most recent put to the same object.

To explain, they described S3's indexing subsystem, which holds all object metadata such as names, tags, and creation times. Every get, put, list, head, or delete goes through the index. More requests hit the index than the storage layer, because head and list requests are served entirely from metadata. At the center of the index is its own storage system, which must be sized to meet S3's availability and durability promises.

Index data is stored across replicas using a quorum-based algorithm, which Tomsen Bukovec called very forgiving of failures. Replicas run on servers in separate availability zones so that data is not correlated on a single fault domain. The failure of one disk, server, rack, or zone affects only a subset of data and never all or a majority of the data for a single object. S3 also caches heavily at the front end. In the eventually consistent design, a read could be routed at random to a cache entry that did not yet reflect the latest write. Quorum guaranteed that reads and writes overlapped at the index storage layer, but the cache did not provide that guarantee, because it was optimized for availability.

Strong consistency: the replicated journal and cache coherency

S3 later became strongly consistent. Tomsen Bukovec said this took time to build and called it the "secret sauce" of S3, which AWS rarely discusses. The requirement was to deliver strong consistency without compromising availability. The team would no longer trade one against the other, and that constraint required a new data structure.

That structure is a replicated journal: a distributed data structure that chains nodes together. A write flows through the storage nodes sequentially, with each node forwarding to the next. When a storage node is written, it learns the sequence number of the value along with the value itself. On a later read, for example through the cache, the sequence number can be retrieved and checked. Tomsen Bukovec called the replicated journal "the heart" of S3's strongly consistent, highly available design.

The host asked about failures, since a sequential chain seems more fragile than an eventually consistent system. Tomsen Bukovec said the second part of the design is a new cache coherency protocol. It keeps the property that multiple servers can receive requests and some are allowed to fail. Tomsen Bukovec calls this a "failure allowance." The replicated journal and the coherency protocol together produce strong consistency.

This came with a real cost in hardware. Tomsen Bukovec recalled a debate in the room with S3 engineers about whether to pass that cost to customers, and said they explicitly decided not to. Strong consistency would be free, would apply to every request, and would not be limited to a particular bucket type. The reasoning was that it should become part of the building block and customers should not have to think about its cost. The host said they found this remarkable, because strong consistency normally adds latency or cost, and they described the latency as unchanged. That latency claim was the host's; Tomsen Bukovec did not discuss latency figures.

Proving correctness with automated reasoning

Tomsen Bukovec stressed that it is one thing to say S3 is strongly consistent on every request and another to know it. S3 runs every kind of workload, and part of its value is that scale decorrelates those workloads. How can the team be sure the model holds everywhere?

Their answer was automated reasoning, which they described as what you would get if computer science and math "got married and had kids." The host asked whether this meant formal methods, and Tomsen Bukovec confirmed it. The team built a proof of the consistency model and runs it on code check-ins to the index subsystem, covering both the caching and storage sublayers. Anyone who changes consistency-related code paths is checked against the proof, so a regression in the consistency model would be caught.

The host asked what such a proof looks like in practice. Tomsen Bukovec answered at a high level. For consistency, the proof covers all the combinations of cases to show the model is correct. S3 also uses formal methods for cross-region replication, to prove that data replicated from one region arrived in another, and to prove the correctness of APIs. They named correctness as a design principle as important as durability, availability, and cost. The goal is to verify correctness on every check-in and every request, not just once. In Tomsen Bukovec's words, "at a certain scale math has to save you," because no one can test every edge case at S3 scale. They offered research papers on the subject, which are linked in the episode notes. The host remarked that formal methods are still uncommon even at infrastructure startups.

Durability: auditors and constant failure

The host raised S3's durability promise, which they described as eleven nines. They noted that even four nines of availability is considered hard in backend systems. With 500 trillion objects, they asked, how do you verify durability on real data rather than relying only on a proof that assumes hardware failure rates?

Tomsen Bukovec said durability is handled mostly in the storage layer and depends on both software and the physical placement of data across disks, servers, racks, availability zones, and regions. Availability zones are physically separate locations, sometimes far apart, and some regions have more than three, giving additional fault domains. The most important element, in their view, is the auditors. More than 200 microservices sit behind a single S3 regional endpoint, each doing one or two things well, loosely coupled through well-defined interfaces. They include health checks, repair systems, and auditor systems. The auditors inspect every byte across the fleet, and when they find something that needs repair, repair systems take over. A significant share of these microservices is dedicated to durability.

The host asked whether someone at S3 can say at any moment what durability was over the past week, month, or year. Tomsen Bukovec answered yes. They also said servers were failing during the conversation itself, because servers always fail. S3 is built on that assumption. Systems constantly assess where a failure affected a node, which bytes were involved, and what repair to start. This all happens separately from the gets and puts customers see; Tomsen Bukovec called it the "whole universe under the hood" of managing bytes at scale. The host contrasted this with their own side project, where a full disk was an unusual event that happened once in three years.

Correlated failure, crash consistency, and failure allowances

Tomsen Bukovec said the key is to think about correlated failure: "if you're thinking about availability at any scale, it's the correlated failure that'll get you." Quorum can tolerate one node failing. If all the nodes are in the same availability zone or rack and fail together, availability suffers and the failure allowance is gone. S3 therefore designs around how workloads are exposed to different levels of failure. When an object is uploaded, it is replicated many times. That replication supports durability, and Tomsen Bukovec emphasized that it supports availability too: if a rack, server, or entire availability zone fails, a copy is still available elsewhere.

They also described crash consistency: a system should always return to a consistent state after a fail-stop failure. If engineers reason about the set of states a system can reach when failures occur, and always assume failures will occur, they can design the microservices to preserve consistency and availability together. Tomsen Bukovec said correlated failures, crash consistency, and failure allowances in caches are the everyday work of S3 engineers.

The host asked how failure allowances differ from the error budgets many companies use loosely. Tomsen Bukovec said a failure allowance is necessary, because assuming no failure leads to "a very bad day for your customer." Using the cache as an example, the allowance is managed through sizing so that customers never notice it. Many microservices exist only to track metrics, and cache sizing is based on those metrics and the size of the underlying system. Tomsen Bukovec argued that because S3's layers are so large and already manage correlated failures and failure allowances, every application built on S3 benefits from them.

Culture: respecting the past while being technically fearless

The host quoted distinguished engineer Andy Warfield. Warfield said he once believed large-scale software was basically code, but learned on S3 that code is inseparable from organizational memory, operational practices, and the scale of the system. The host asked how engineers manage such an intimidating system.

Tomsen Bukovec pointed to culture and commitment. S3 engineers range from people just out of school to people who have been on the team for 15 years. Two Amazon engineering tenets pull against each other. "Respect what came before" says that anything that has worked for many years deserves respect, which argues for conservatism. "Be technically fearless" pushes for invention. Tomsen Bukovec said the tension is part of what makes the work enjoyable. New capabilities must preserve S3's existing properties: it has to keep working, with the same durability and availability. At the same time, work such as conditionals, native Iceberg support, and vectors extends the foundation for future applications. In Tomsen Bukovec's view, the team embodies both tenets every day. In the outro, the host said this pair of tenets was one of their favorite parts of the conversation. They argued that a system this critical could easily become purely conservative and fall behind.

From Hive to Iceberg, and SQL as the interface

The host asked whether S3 is "done," since it can already store any blob. Tomsen Bukovec returned to the rise of Parquet around 2020. In their account, Hive gave Hadoop file-system-style access to S3's unstructured storage. Iceberg replaced Hive by giving tabular access to Parquet data, including compaction and table maintenance. Tomsen Bukovec said they believe the world's tabular data will live in S3 in the future.

As an example, they cited Supabase's announcement the week before. According to Tomsen Bukovec, Supabase's Postgres database would do secondary writes directly into S3 Tables, and its Postgres vector extension would integrate with S3 Vectors. Tomsen Bukovec argued that SQL is the lingua franca of data and that LLMs have been trained on decades of SQL (the host added Python). Many AWS customers already know the S3 API, but S3 Tables let anyone who knows SQL work with data in S3 without knowing S3 or cloud development, whether that user is a human or an AI agent. Tomsen Bukovec predicted this would grow quickly in the coming years.

S3 Vectors: why vectors belong in storage

The host asked what it takes to build a new data primitive like vectors. Tomsen Bukovec compared the situation to tabular data. People once put tables into databases mainly to query them, even when they did not really need a database. Open formats like Parquet later let that data live in S3. Tomsen Bukovec described S3 Vectors as doing the same for vectors, which today often live in dedicated vector databases.

They then described what they called one of the great ironies of data: "you have to know your data to know your data." You need to know the schema, the types, and where the data is. As data lakes become data oceans, this gets harder. Embedding models understand the data for you, and their output is a vector. Tomsen Bukovec said a company's knowledge is not organized in rows and columns. It lives in PDFs, phones, customer-care audio recordings that capture how customers feel, whiteboards, and documents spread across dozens of systems. Understanding what data you have across those formats is a real problem that AI models can help with. They said these models have improved greatly in the last 18 to 24 months. What customers lacked was a place to store billions of vectors, and S3 Vectors was built for that. Tomsen Bukovec emphasized that it is not a database: it has S3's cost structure and scale, applied to vector storage.

How S3 Vectors works

The host asked whether vectors are built on existing object storage. Tomsen Bukovec said no. S3 Tables build on objects, because Parquet files are objects, but vectors required a new data structure and data type.

The core difficulty is nearest-neighbor search in high-dimensional space. A naive approach compares a query against every vector, which is very expensive. S3 does not keep vectors in memory; they are stored across S3's fleet, yet queries still need low latency. At launch, Tomsen Bukovec said, warm queries were returning in about 100 milliseconds or less. They called this "not database fast, but pretty fast."

The approach is to precompute what Tomsen Bukovec called "vector neighborhoods": clusters of similar vectors, such as vectors for one type of dog. These are computed offline and asynchronously so they do not affect query performance. When a new vector is inserted, it is added to one or more neighborhoods. At query time, S3 first performs a much smaller search to find the nearest neighborhoods. Only those vectors are loaded from S3 into fast memory, where the nearest-neighbor algorithm runs. Tomsen Bukovec gave the following limits: up to 2 billion vectors per index and up to 20 trillion vectors per vector bucket, with warm queries at 100 milliseconds or less.

Tomsen Bukovec connected this to an S3 service tenet: "scale is to your advantage." A design cannot get worse as it grows; it has to get better. They cited S3 itself as the example: the bigger it gets, the more decorrelated the workloads running on it become. For vectors, the team asked how to make 100 milliseconds "just the start" and how to ensure that S3's characteristics improve as more vectors are stored.

The 50 TB object limit and how the roadmap is set

The host asked why the largest object size is 50 terabytes. Tomsen Bukovec pointed out that the limit is ten times the original 5 TB. When customers ask what could possibly be that large, the answer is high-resolution video. Size limits reflect optimizing the underlying systems for particular patterns. Raising the limit tenfold, like the tenfold increase in batch operations scale announced the week before, meant optimizing for the new typical distribution of work. Tomsen Bukovec said S3 has few limits and will keep adjusting them as workload distributions change. Larger objects are appearing because of better cameras and phones, and the team wanted customers to be able to grow without restriction.

On the roadmap, Tomsen Bukovec said S3 has launched over a thousand capabilities since 2020. About 90% of the roadmap comes from explicit customer requests, such as larger objects for media customers or batch operations improvements. Some capabilities are invented by watching what customers do with their data, and vectors fall into that category. The team asked how to make data usable in an industry-standard way, like Iceberg for tabular data, and how to make it usable now that embeddings can provide semantic understanding, if only storing billions of vectors were affordable. S3's goal is to remove two constraints: the cost of data and the difficulty of working with it. When both are addressed, Tomsen Bukovec said, the team has what they call a "product shape." They described S3 as a living, breathing organism whose shape changes while staying consistent with the traits people expect. The host compared it to a plant, and Tomsen Bukovec mentioned being a former Peace Corps forestry volunteer who often uses metaphors from nature.

Simplicity as a discipline

The host asked what makes engineering at this scale so different from engineering at a startup. Tomsen Bukovec's answer was simplification. S3 is a very complex system, so each microservice must do one or two things well; otherwise a distributed system becomes unmaintainable over time. Simplicity applies in two places. In the user model, it means a simple API, and now SQL through S3 Tables and semantic understanding through embeddings, instead of annotating a whole metadata layer by hand. Under the hood, Tomsen Bukovec said, S3 engineering meetings constantly return to implementing each capability as simply as possible.

Who works on S3

S3 hires engineers at all career stages. The common trait Tomsen Bukovec emphasized was ownership: a personal commitment to preserving each customer's bytes and keeping them useful, so customers can focus on their applications rather than storage. Tomsen Bukovec argued that every modern business is a data business, and that the data teams feel responsible for that data.

For mid-career engineers who want to work on deep infrastructure, Tomsen Bukovec recommended "relentless curiosity." Working on a system that keeps redefining storage means you are not coloring within the lines. You draw today's lines knowing you may have to erase and redraw them. They said they tell their own three children, who are at university or graduate school, the same thing: step back and read the latest research. They pointed to the papers they would share, such as bringing formal methods into storage systems or thinking about failure in new ways, as examples of that creativity. The host added that startups such as Turbopuffer are now building on S3 as a base layer. Tomsen Bukovec said it was exciting to see so many kinds of infrastructure built on S3.

Closing: multimodal embeddings

Asked for a paper or book recommendation, Tomsen Bukovec pointed to research on multimodal embedding models. They argued that the world we experience is multimodal, so our understanding of data should be too. In their view, the next generation of data lakes will be built on metadata and semantic understanding, which will be created through vectors and searched across multiple modalities. They expect vectors to become very large, especially at the price point AWS introduced for S3 Vectors, and said "we're just getting started" with understanding our data. Their book recommendation was outside computer science: a book on ecology and supporting native bees and insects. The title was not given in the conversation.