Roll Forward, Toggle Often: Robert Erez on How Teams Actually Ship Software

Open on YouTube ↗
Overview

Robert Erez is a principal engineer at Octopus Deploy and has worked on CI/CD and deployment tooling for more than a decade. On The Pragmatic Engineer, he and the host, a former teammate from Skype, discussed a practical question: what does safe, efficient software delivery look like across many different organizations? Erez consistently takes a pragmatic line. Conference talks tend to present one right way to deploy, but his customers mostly "just want to ship software." Across Kubernetes, GitOps, platform teams, AI, progressive delivery and self-hosting, he argues that context matters more than dogma, and that feature toggles and rolling forward are more reliable than rollbacks.

25 min read

A canary called New Zealand

Erez's interest in delivery started at Skype, where he and the host worked on the web team. The team's product included the Outlook.com plugin, which the host recalled had something like 400 million monthly users. The official process was a weekly release that required sign-off from a change advisory board. Erez found this odd. The software ran on the web, on Azure, and the team had full access to push whenever it wanted, yet changes made during the week were held back.

With the support of their manager, the team worked around the system. Code was committed, went through several layers of testing, moved to staging, and then shipped to production during the week. Erez describes this as a form of continuous delivery that he was proud of. The team also ran canary deployments, releasing a new build to a small share of customers first, and that share was always New Zealand.

He gave several reasons. New Zealand is the first country of meaningful size to reach a new date, so it was naturally first in a time-based rollout. People there speak English, which made bug reports easy to understand. And, he admitted with an apology to New Zealanders, the market was small enough that nobody minded much if a bug shipped and had to be fixed quickly. For Erez, the experience showed how continuous delivery techniques could beat big-bang releases, and it shaped his understanding of progressive delivery.

Joining Octopus Deploy as an early engineer

After a few years at Skype, Erez and his wife moved back to Australia. He briefly worked at a friend's company that happened to use Octopus Deploy, a deployment tool originally built in Brisbane. When he learned Octopus was hiring, he applied because he liked the CI/CD space and thought it had interesting problems.

He believes he was roughly employee number eight or nine. He calls it a startup, though "not startup in the sense of Silicon Valley wild parties": everyone, including the CEO, was an engineer. Ideas were discussed and shipped quickly, and the engineers also handled marketing and support. The company has grown considerably since.

The maturity ladder: from YOLO to continuous deployment

The host asked why Octopus focused on deployment, and whether that is the same as continuous delivery. Erez described a maturity progression. The first stage, which the host called "YOLO," is deploying straight to production or SSHing into servers. Erez acknowledged that most engineers have worked somewhere like that.

The second stage is continuous integration: merging changes into a single branch continually and running tests against it. The third is continuous delivery. Beyond unit and integration tests, you also test the deployment process itself, so that clicking "deploy" at any point would successfully reach production. The fourth, which not every company reaches, is continuous deployment, where changes flow to production automatically without intervention.

The key difference between delivery and deployment is whether changes reach production without manual steps. With continuous delivery, changes move through environments such as dev, test and staging. Some of those steps may still be manual, for example refreshing a test environment only weekly for testers, but the pipeline could run end to end automatically if the team wanted it to.

Erez does not think every team should go all the way to continuous deployment. Regulated industries may still need review boards and releases scheduled for times when the right people are available. If changes are continually passing through testing and being promoted through environments, then pressing the production button only once a week is "fine." Most of the hard work is done and risk has been mitigated. In his framing, the point of the whole process is to feel pain as early as possible and de-risk everything before the final step.

Why Kubernetes won

The host was surprised that Octopus's site prominently mentions Kubernetes. Erez called Kubernetes "the platform of the moment." He traced it to Google's internal Borg system and said, while noting he could not read Google's mind, that releasing Kubernetes partly helped level the playing field with other cloud vendors. The host mentioned that Kat Cosgrove had made a similar point on the podcast: Kubernetes made it easier to move workloads between AWS and Google Cloud, so choosing a provider other than AWS became less risky.

Kubernetes arrived when several container orchestration options were competing, including Nomad, Docker Swarm and, as the host added, CoreOS's Fleet. Erez also mentioned Azure Service Fabric as a non-container attempt at the same problem of hosting and orchestrating many microservices. In his view, Kubernetes won because its mechanics appealed to engineers and DevOps teams, and because cloud vendors adopted it, which made it easy to use across platforms.

Kubernetes, far from the cloud

Erez found it "kind of funny" that Kubernetes is marketed as cloud native, since many Octopus customers run it on premises. A significant number run it on their own VMs and server farms, or on cloud VMs where they maintain Kubernetes themselves for more control. He says this is especially common in financial services. There, organizations want end-to-end control of their infrastructure but still want the declarative model that lets application and ops teams define what runs and how.

He placed Kubernetes alongside other declarative tools such as Terraform and Puppet. You describe the desired state, and the system keeps reality matching it. If you ask for three replicas and a pod dies, Kubernetes starts another.

When asked what "on premise" means in practice, Erez said it covers everything from co-located data centers to machines in a closet. Some customers run small Kubernetes clusters in point-of-sale systems across hundreds of stores, each cluster independent. At the scale of thousands of clusters that all follow GitOps practices and pull state from a Git repository, the repository becomes the bottleneck or requests get throttled, so these customers have to find workarounds.

His most striking example came from a conversation at KubeCon with a customer running Kubernetes clusters on research vessels, meaning actual ships at sea. He didn't learn what the ships do. The deployment problem, though, was clear. A ship may be away for weeks or months, so it cannot receive updates when a deployment happens. It has to pick them up when it returns to port, and Erez and the customer were discussing how that process could work.

What GitOps actually is

Erez named GitOps as one of the biggest trends he sees. The idea extends Kubernetes' continuous reconciliation back into Git: engineers change a declarative definition in a repository, and a process makes sure the cluster matches it. He said Weaveworks coined the term around 2017. It grew alongside Kubernetes and was formalized in the early 2020s into four pillars:

  1. Declarative. You define the desired state of your infrastructure, rather than an imperative sequence of steps whose result is harder to predict.
  2. Versioned and immutable. The desired state lives somewhere you can point to, such as a commit SHA or tag, and that reference shouldn't change. This also makes auditing easier.
  3. Pulled, not pushed. An agent pulls the state from its source and applies it to the cluster.
  4. Continuously reconciled. If someone deletes a pod or deployment by hand, the system repairs the drift.

The host asked whether Git really provides immutability, since history can be rewritten. Erez agreed. Depending on how the GitOps agent is configured, especially if it points at tags, which can be moved, the state can change. That is why best practices discourage relying on tags.

His main point is that none of the four pillars mentions Git. He thinks the name leads people to assume everything must live in Git. Secrets are the obvious counterexample. The community has built workarounds such as sealed secrets, which encrypt secrets so they can be committed; the host called that "a terrible idea." For Erez, these workarounds show that some things simply don't need to be in Git, as long as you keep control over versioning and immutability.

Where the GitOps mindset breaks down

Erez sees GitOps spreading mainly because Kubernetes is spreading in enterprises. When organizations adopt Kubernetes and ask how to deploy to it, GitOps becomes the de facto answer. There are also some projects and experiments applying it beyond Kubernetes, for example continuous reconciliation of Terraform state, but Kubernetes remains its core territory.

He warned against absolutism. Teams hear at conferences and in blog posts that GitOps is the right approach, but real deployments involve steps that don't fit a declarative, everything-in-Git model: running smoke tests, sending notifications, performing database updates. Tools such as Argo Workflows and Argo Rollouts exist to "GitOpsify" these steps, and that works for some teams. But Erez believes some people treat Git as a hammer and everything else as a nail.

When he talks through these steps with customers, he says, they recognize that the goal is simply to ship software. If GitOps works end to end, good. If it covers part of the process and imperative mechanisms handle the rest, that is also fine. There are tens of thousands of companies delivering software, and most of them are not at conferences.

The rise of platform teams

Erez sees platform teams as a newer standard structure that grew out of DevOps. In the older model, developers threw code over the wall to ops. DevOps gave engineering teams ownership of operations, so they felt the pain and fixed problems faster. As organizations grew, two problems appeared. Separate "DevOps teams" emerged, recreating the split DevOps was meant to remove. And application teams each built deployment processes from scratch, so processes diverged, people couldn't easily move between teams, and developers suffered context overload from keeping up with cloud best practices.

The host recalled a five-person mobile team where one person effectively spent half their time on Jenkins configuration, and people had to draw straws for the job because everyone wanted to build features. Erez agreed that nobody enjoys spending their time managing pipelines.

Platform teams, as he describes them, don't own the whole process the way an ops team did. They define best practices and provide self-service tooling, often through an internal developer portal where application teams can, for example, create a new project from a template. Operational ownership stays with the application teams, so they keep DevOps' close feedback loop without needing expertise in every deployment method. Erez stressed that not every company needs a platform team. Small companies may do fine with application teams practicing DevOps, but platform teams bring "sanity and control and focus" to larger organizations.

AI and CI/CD: less about speed, more about risk

On AI, "the elephant in the room," Erez said it is still very early. He expects CI/CD's changes to lag behind how development teams adopt AI, and Octopus is currently talking to customers to learn how they use it. He noted, half-jokingly, that Octopus was one of the few companies at KubeCon without AI "plastered all over" its booth. The company does ship AI features, including an MCP server and a recovery agent that can review logs and tasks, but it wants to add them only where they are actually useful.

His main prediction is more velocity: much more code coming through pipelines. That raises the question of what pipelines should optimize for. Today, much effort goes into shortening feedback loops so waiting engineers don't have to context switch. If AI writes most of the code, Erez suggested, that may matter less. If the engineer has moved on and an agent can watch the build, review failures and issue fixes, it may not matter much whether the pipeline takes 30 minutes or 20. He expects the emphasis to shift from pipeline speed toward reducing the risk that comes from AI-generated code, though he said exactly what that looks like "remains to be seen."

He expects progressive delivery, and feature toggles in particular, to become much more common. Toggles let teams ship code as fast as they want while controlling the rollout separately, which decouples deployment from release. In a world with more AI-generated code, he thinks agents being able to use toggles to react quickly will matter more than it does today.

Progressive delivery: canaries, blue/green, and toggles

Erez calls progressive delivery the next step beyond continuous delivery. Instead of sending a change to an environment all at once, you release it in a controlled way. In a canary deployment, you run version two beside version one and usually use a network traffic manager to send a percentage of traffic to the new version, raising or lowering it based on mature observability. The name comes from canaries in coal mines, whose reaction to toxic gases gave miners an early warning to get out.

In blue/green deployments, the new version runs alongside the old one while the old one still takes all traffic. You can test it directly, perhaps warm it up to avoid cold starts, and then switch all traffic over. Erez described it as a canary that goes straight to 100 percent, with validation done before customers reach it.

He considers feature toggles (or feature flags, which he uses interchangeably) the most useful strategy, particularly for application delivery. A toggle is a variable, usually tied to an external service, whose value selects a code path. He gave several reasons for preferring it to canaries:

  • Granularity. A toggle can change a single line of behavior while everything else stays the same. With a canary, the unit of change is the whole app, so 20 commits since the last release are all tested at once.
  • Targeting. Toggles support complex rules, such as everyone in Germany with a particular product in their basket. That is very hard to express with network traffic rules.
  • Reversal speed. Undoing a canary may mean redeploying the old version, which can take minutes or more. Turning off a toggle takes seconds.
  • Timing. With versioned deployments, a feature goes live when the deployment happens, so teams, perhaps ten at once, have to watch logs at that moment. With toggles, you might deploy on Monday and release on Tuesday when you are ready and watching.

Versioned deployments still have a place, he said. They suit infrastructure changes where no application toggle applies, and changes involving schemas.

Schema changes and the SaaS/self-hosted split

Schema changes are "the big problem" in any progressive delivery, according to Erez. Handling them requires teams mature enough to understand the pitfalls and roll changes out gradually over multiple stages. Octopus finds this especially hard because it ships both a SaaS product and a self-hosted version, which Erez calls "the best and worst of both worlds."

In SaaS, Octopus controls which version runs where, so it can stage an expand-and-contract migration and confirm each stage is complete before moving on. Self-hosted customers might jump from version one to version six, skipping the intermediate stages. On the other hand, those customers choose when to upgrade, so Octopus can ask them to take backups first, and they may accept more downtime during migrations than a SaaS product could.

Why "rollback" is the wrong conversation

Erez called rollbacks "always a spicy one." Customers often ask why Octopus has no rollback button. In a fully stateless system, rollback is easy; with GitOps, you can revert a commit and push. But most systems have state, especially databases. Rolling back code can leave it out of step with the database schema. You can sometimes ship a reverse migration alongside a migration, but that isn't always possible, and it doesn't answer what to do with the data.

Octopus's advice to customers is therefore to avoid talking about rollback at all: always roll forward. If version two has a bug, the fix is version three, not version one. This is what hotfix processes and fast feedback loops are for. Normally changes might pass through dev, staging, prod and approval gates. For a serious bug, the safest option may be to hotfix the current version and push it out as quickly as possible, with the build pipeline as the bottleneck and speed depending on the team's appetite for risk. If the failure was in the deployment mechanism itself, recovery can be faster still.

He said customers sometimes claim they roll back all the time. When asked what they do with a schema change, "they kind of stop" and realize it has been sheer luck that they never hit that case.

The host asked whether reverting is fine for stateless application logic. Erez said that ideally such a reversal happens through a feature flag, changing the code path instead of redeploying. Schema issues still need care, but the host noted, and Erez agreed, that the flagged code paths can be designed to stay consistent with whichever schema version is deployed.

Keeping feature flags from becoming a mess

The host raised the downside of heavy flag use, which they had seen at Uber: stale flags scattered across the codebase. Erez agreed emphatically. Octopus uses OpenFeature as its SDK and has built a wrapper so that each toggle in code records its owning team and an expiry date. Nothing breaks when the date passes, but the CI process notifies the team that the toggle may no longer be needed. Erez admitted that if he logged in, he would probably find notifications asking him to remove some of his own.

He compared this to weeding in the gardening metaphor for code: it needs constant attention. Some tools use observability data to show when a toggle was last evaluated, which is a useful signal. He also described a sequencing problem. A toggle lives in two places, the code and the flag platform. You should remove it from the code first, but that removal may take two weeks to reach production, and by then you've forgotten you still need to delete the configuration. Mechanisms that track the removal through to production and then report that the configuration is safe to delete help close that gap.

From shared test environments to ephemeral ones

Asked how development environments are evolving, Erez said there is no single pattern. Dev, test and prod is the most common setup, and even that simplifies things. In CD terms, "dev" is usually the first point of integration, where you check that the deployment works at all. Test is often kept close to production, sometimes with sanitized data, for QA and product teams.

He sees dev environments becoming less useful and ephemeral environments growing. An engineer working on a feature branch wants to confirm it works, let teammates see it, and then switch to something else, which is hard if it only runs on their laptop. An ephemeral environment spins up from the pre-merge branch with the dependencies it needs, deploys the app as if to a real environment, provides a URL others can use, and is torn down when the PR merges. This replaces arrangements like one test environment per tester, or a single shared test environment that testers must coordinate access to.

The host noted that talk about cloud development environments and preview environments seems to have faded, even though the technology exists. Erez said they are easy to discuss for simple cases but get tricky with multiple apps, stateful services and realistic data. The host suggested AI agents make them more valuable, since running the code and seeing it work, especially with a UI, is a strong form of validation. Erez agreed: whenever an agent validates its own work by testing and exploring, it is using an ephemeral environment, even if no human ever sees it.

Running Octopus Cloud

Erez is no longer on the team that runs Octopus's SaaS offering, but he gave some history. When Octopus first tried SaaS, around 2020 by his recollection, every customer got a dedicated VM with the self-hosted product installed. He estimated this cost about $100 per customer per month while customers paid something like $20. It was an experiment to test demand, and he credited the CEO for being willing to fund it. Moving from downloadable software to hosting was a big step for the company.

Demand turned out to be real, so a small group rebuilt the service on Kubernetes. Erez worked on getting Octopus running on Linux and in containers, while others built the dynamic worker infrastructure, partly to "stop losing money." The architecture uses what Octopus calls a "reef," in keeping with its nautical naming. A reef is a cell containing the resources for a set of customer instances, including a cluster and an Azure database, with some resources shared, and each customer instance runs as a pod. Today a dedicated team runs the service for several thousand customers and many thousands of deployments a month.

A current project aims to make deployments more resilient. Deployments run as imperative steps, and much of their state is held in memory, so upgrades require pausing tasks and restarting instances, which means some downtime. The goal is to get as close to zero as possible. Erez noted that cutting downtime from five minutes to one is "just work," while going from ten seconds to zero is a much bigger shift, and he isn't sure they will get there.

Why self-hosting isn't going away

Octopus rolls cloud updates out gradually, instance by instance, over a few days. On-prem adoption is much slower. When Erez looked at the numbers, it took about 200 days for 50 percent of on-prem customers to pick up a change and more than 400 days for 75 percent. Some customers run versions from five to seven years ago, so Octopus has to support upgrades across a range such as 2023.1 to 2026.4, with all the schema-migration baggage that brings.

Asked why Octopus doesn't just drop old versions, Erez said most of its customers are still on-prem: banks, financial institutions and governments that want full control of upgrades and downtime, whether on their own hardware or their own cloud accounts. He doesn't expect that to go away. Octopus has become more willing to deprecate old capabilities in recent years, partly thanks to adopting feature toggles, but he believes self-hosting will stay because for many organizations the cloud isn't viable or doesn't meet compliance requirements.

The host suggested that on-prem support may mean less competition for infrastructure vendors, and Erez agreed. The host also described long-unupgraded customers as potentially loyal ones. Erez related to that from his earlier job: if something critical like a deployment system works, customers leave it alone. That frustrates Octopus engineers who want to ship new features, but he praised the support team for helping customers on old versions. The host compared this to LLMs: when a new model version breaks a tuned workflow, some organizations may pay to pin a version or run it themselves.

Getting started, and what to read

For engineers wanting to move into progressive delivery, Erez's advice was to start with one feature toggle. Flipping something in production can feel scary compared with relying on tests well before code runs. But he said it becomes addictive, which is exactly why flag hygiene matters. He has shipped bugs behind toggles himself. The difference, he said, is that instead of panicking at 2 a.m. on call and wondering whether to build a new version or force a redeploy, you can turn the feature off, stop the bleeding, and then investigate calmly. Once you experience that, he said, you want to use toggles for everything.

His book recommendations were The Phoenix Project by Gene Kim, which their Skype manager gave everyone. Parts may be dated, he said, but its central idea of engineers owning the operational side of what they ship underpins everything discussed in the episode. He also recommended Radical Candor by Kim Scott for communicating directly while showing care, which he admits he needs as an engineer who can be blunt. For fun, he recommended anything by the Australian hard science fiction author Greg Egan, naming Diaspora and Schismatrix. He described Egan as a mathematician who builds entire stories from one altered physical premise, such as a speed of light that isn't absolute, and works out the consequences.