Roll Forward, Toggle Often: Robert Erez on How Teams Actually Ship Software
The Pragmatic EngineerRobert Erez is a principal engineer at Octopus Deploy and has worked on CI/CD and deployment tooling for more than a decade. On The Pragmatic Engineer, he and the host, a former teammate from Skype, discussed a practical question: what does safe, efficient software delivery look like across many different organizations? Erez consistently takes a pragmatic line. Conference talks tend to present one right way to deploy, but his customers mostly "just want to ship software." Across Kubernetes, GitOps, platform teams, AI, progressive delivery and self-hosting, he argues that context matters more than dogma, and that feature toggles and rolling forward are more reliable than rollbacks.
A canary called New Zealand
Erez's interest in delivery started at Skype, where he and the host worked on the web team. The team's product included the Outlook.com plugin, which the host recalled had something like 400 million monthly users. The official process was a weekly release that required sign-off from a change advisory board. Erez found this odd. The software ran on the web, on Azure, and the team had full access to push whenever it wanted, yet changes made during the week were held back.
With the support of their manager, the team worked around the system. Code was committed, went through several layers of testing, moved to staging, and then shipped to production during the week. Erez describes this as a form of continuous delivery that he was proud of. The team also ran canary deployments, releasing a new build to a small share of customers first, and that share was always New Zealand.
He gave several reasons. New Zealand is the first country of meaningful size to reach a new date, so it was naturally first in a time-based rollout. People there speak English, which made bug reports easy to understand. And, he admitted with an apology to New Zealanders, the market was small enough that nobody minded much if a bug shipped and had to be fixed quickly. For Erez, the experience showed how continuous delivery techniques could beat big-bang releases, and it shaped his understanding of progressive delivery.
Joining Octopus Deploy as an early engineer
After a few years at Skype, Erez and his wife moved back to Australia. He briefly worked at a friend's company that happened to use Octopus Deploy, a deployment tool originally built in Brisbane. When he learned Octopus was hiring, he applied because he liked the CI/CD space and thought it had interesting problems.
He believes he was roughly employee number eight or nine. He calls it a startup, though "not startup in the sense of Silicon Valley wild parties": everyone, including the CEO, was an engineer. Ideas were discussed and shipped quickly, and the engineers also handled marketing and support. The company has grown considerably since.
The maturity ladder: from YOLO to continuous deployment
The host asked why Octopus focused on deployment, and whether that is the same as continuous delivery. Erez described a maturity progression. The first stage, which the host called "YOLO," is deploying straight to production or SSHing into servers. Erez acknowledged that most engineers have worked somewhere like that.
The second stage is continuous integration: merging changes into a single branch continually and running tests against it. The third is continuous delivery. Beyond unit and integration tests, you also test the deployment process itself, so that clicking "deploy" at any point would successfully reach production. The fourth, which not every company reaches, is continuous deployment, where changes flow to production automatically without intervention.
The key difference between delivery and deployment is whether changes reach production without manual steps. With continuous delivery, changes move through environments such as dev, test and staging. Some of those steps may still be manual, for example refreshing a test environment only weekly for testers, but the pipeline could run end to end automatically if the team wanted it to.
Erez does not think every team should go all the way to continuous deployment. Regulated industries may still need review boards and releases scheduled for times when the right people are available. If changes are continually passing through testing and being promoted through environments, then pressing the production button only once a week is "fine." Most of the hard work is done and risk has been mitigated. In his framing, the point of the whole process is to feel pain as early as possible and de-risk everything before the final step.
Why Kubernetes won
The host was surprised that Octopus's site prominently mentions Kubernetes. Erez called Kubernetes "the platform of the moment." He traced it to Google's internal Borg system and said, while noting he could not read Google's mind, that releasing Kubernetes partly helped level the playing field with other cloud vendors. The host mentioned that Kat Cosgrove had made a similar point on the podcast: Kubernetes made it easier to move workloads between AWS and Google Cloud, so choosing a provider other than AWS became less risky.
Kubernetes arrived when several container orchestration options were competing, including Nomad, Docker Swarm and, as the host added, CoreOS's Fleet. Erez also mentioned Azure Service Fabric as a non-container attempt at the same problem of hosting and orchestrating many microservices. In his view, Kubernetes won because its mechanics appealed to engineers and DevOps teams, and because cloud vendors adopted it, which made it easy to use across platforms.
Kubernetes, far from the cloud
Erez found it "kind of funny" that Kubernetes is marketed as cloud native, since many Octopus customers run it on premises. A significant number run it on their own VMs and server farms, or on cloud VMs where they maintain Kubernetes themselves for more control. He says this is especially common in financial services. There, organizations want end-to-end control of their infrastructure but still want the declarative model that lets application and ops teams define what runs and how.
He placed Kubernetes alongside other declarative tools such as Terraform and Puppet. You describe the desired state, and the system keeps reality matching it. If you ask for three replicas and a pod dies, Kubernetes starts another.
When asked what "on premise" means in practice, Erez said it covers everything from co-located data centers to machines in a closet. Some customers run small Kubernetes clusters in point-of-sale systems across hundreds of stores, each cluster independent. At the scale of thousands of clusters that all follow GitOps practices and pull state from a Git repository, the repository becomes the bottleneck or requests get throttled, so these customers have to find workarounds.
His most striking example came from a conversation at KubeCon with a customer running Kubernetes clusters on research vessels, meaning actual ships at sea. He didn't learn what the ships do. The deployment problem, though, was clear. A ship may be away for weeks or months, so it cannot receive updates when a deployment happens. It has to pick them up when it returns to port, and Erez and the customer were discussing how that process could work.
What GitOps actually is
Erez named GitOps as one of the biggest trends he sees. The idea extends Kubernetes' continuous reconciliation back into Git: engineers change a declarative definition in a repository, and a process makes sure the cluster matches it. He said Weaveworks coined the term around 2017. It grew alongside Kubernetes and was formalized in the early 2020s into four pillars:
- Declarative. You define the desired state of your infrastructure, rather than an imperative sequence of steps whose result is harder to predict.
- Versioned and immutable. The desired state lives somewhere you can point to, such as a commit SHA or tag, and that reference shouldn't change. This also makes auditing easier.
- Pulled, not pushed. An agent pulls the state from its source and applies it to the cluster.
- Continuously reconciled. If someone deletes a pod or deployment by hand, the system repairs the drift.
The host asked whether Git really provides immutability, since history can be rewritten. Erez agreed. Depending on how the GitOps agent is configured, especially if it points at tags, which can be moved, the state can change. That is why best practices discourage relying on tags.
His main point is that none of the four pillars mentions Git. He thinks the name leads people to assume everything must live in Git. Secrets are the obvious counterexample. The community has built workarounds such as sealed secrets, which encrypt secrets so they can be committed; the host called that "a terrible idea." For Erez, these workarounds show that some things simply don't need to be in Git, as long as you keep control over versioning and immutability.
Where the GitOps mindset breaks down
Erez sees GitOps spreading mainly because Kubernetes is spreading in enterprises. When organizations adopt Kubernetes and ask how to deploy to it, GitOps becomes the de facto answer. There are also some projects and experiments applying it beyond Kubernetes, for example continuous reconciliation of Terraform state, but Kubernetes remains its core territory.
He warned against absolutism. Teams hear at conferences and in blog posts that GitOps is the right approach, but real deployments involve steps that don't fit a declarative, everything-in-Git model: running smoke tests, sending notifications, performing database updates. Tools such as Argo Workflows and Argo Rollouts exist to "GitOpsify" these steps, and that works for some teams. But Erez believes some people treat Git as a hammer and everything else as a nail.
When he talks through these steps with customers, he says, they recognize that the goal is simply to ship software. If GitOps works end to end, good. If it covers part of the process and imperative mechanisms handle the rest, that is also fine. There are tens of thousands of companies delivering software, and most of them are not at conferences.
The rise of platform teams
Erez sees platform teams as a newer standard structure that grew out of DevOps. In the older model, developers threw code over the wall to ops. DevOps gave engineering teams ownership of operations, so they felt the pain and fixed problems faster. As organizations grew, two problems appeared. Separate "DevOps teams" emerged, recreating the split DevOps was meant to remove. And application teams each built deployment processes from scratch, so processes diverged, people couldn't easily move between teams, and developers suffered context overload from keeping up with cloud best practices.
The host recalled a five-person mobile team where one person effectively spent half their time on Jenkins configuration, and people had to draw straws for the job because everyone wanted to build features. Erez agreed that nobody enjoys spending their time managing pipelines.
Platform teams, as he describes them, don't own the whole process the way an ops team did. They define best practices and provide self-service tooling, often through an internal developer portal where application teams can, for example, create a new project from a template. Operational ownership stays with the application teams, so they keep DevOps' close feedback loop without needing expertise in every deployment method. Erez stressed that not every company needs a platform team. Small companies may do fine with application teams practicing DevOps, but platform teams bring "sanity and control and focus" to larger organizations.
AI and CI/CD: less about speed, more about risk
On AI, "the elephant in the room," Erez said it is still very early. He expects CI/CD's changes to lag behind how development teams adopt AI, and Octopus is currently talking to customers to learn how they use it. He noted, half-jokingly, that Octopus was one of the few companies at KubeCon without AI "plastered all over" its booth. The company does ship AI features, including an MCP server and a recovery agent that can review logs and tasks, but it wants to add them only where they are actually useful.
His main prediction is more velocity: much more code coming through pipelines. That raises the question of what pipelines should optimize for. Today, much effort goes into shortening feedback loops so waiting engineers don't have to context switch. If AI writes most of the code, Erez suggested, that may matter less. If the engineer has moved on and an agent can watch the build, review failures and issue fixes, it may not matter much whether the pipeline takes 30 minutes or 20. He expects the emphasis to shift from pipeline speed toward reducing the risk that comes from AI-generated code, though he said exactly what that looks like "remains to be seen."
He expects progressive delivery, and feature toggles in particular, to become much more common. Toggles let teams ship code as fast as they want while controlling the rollout separately, which decouples deployment from release. In a world with more AI-generated code, he thinks agents being able to use toggles to react quickly will matter more than it does today.
Progressive delivery: canaries, blue/green, and toggles
Erez calls progressive delivery the next step beyond continuous delivery. Instead of sending a change to an environment all at once, you release it in a controlled way. In a canary deployment, you run version two beside version one and usually use a network traffic manager to send a percentage of traffic to the new version, raising or lowering it based on mature observability. The name comes from canaries in coal mines, whose reaction to toxic gases gave miners an early warning to get out.
In blue/green deployments, the new version runs alongside the old one while the old one still takes all traffic. You can test it directly, perhaps warm it up to avoid cold starts, and then switch all traffic over. Erez described it as a canary that goes straight to 100 percent, with validation done before customers reach it.
He considers feature toggles (or feature flags, which he uses interchangeably) the most useful strategy, particularly for application delivery. A toggle is a variable, usually tied to an external service, whose value selects a code path. He gave several reasons for preferring it to canaries:
- Granularity. A toggle can change a single line of behavior while everything else stays the same. With a canary, the unit of change is the whole app, so 20 commits since the last release are all tested at once.
- Targeting. Toggles support complex rules, such as everyone in Germany with a particular product in their basket. That is very hard to express with network traffic rules.
- Reversal speed. Undoing a canary may mean redeploying the old version, which can take minutes or more. Turning off a toggle takes seconds.
- Timing. With versioned deployments, a feature goes live when the deployment happens, so teams, perhaps ten at once, have to watch logs at that moment. With toggles, you might deploy on Monday and release on Tuesday when you are ready and watching.
Versioned deployments still have a place, he said. They suit infrastructure changes where no application toggle applies, and changes involving schemas.
Schema changes and the SaaS/self-hosted split
Schema changes are "the big problem" in any progressive delivery, according to Erez. Handling them requires teams mature enough to understand the pitfalls and roll changes out gradually over multiple stages. Octopus finds this especially hard because it ships both a SaaS product and a self-hosted version, which Erez calls "the best and worst of both worlds."
In SaaS, Octopus controls which version runs where, so it can stage an expand-and-contract migration and confirm each stage is complete before moving on. Self-hosted customers might jump from version one to version six, skipping the intermediate stages. On the other hand, those customers choose when to upgrade, so Octopus can ask them to take backups first, and they may accept more downtime during migrations than a SaaS product could.
Why "rollback" is the wrong conversation
Erez called rollbacks "always a spicy one." Customers often ask why Octopus has no rollback button. In a fully stateless system, rollback is easy; with GitOps, you can revert a commit and push. But most systems have state, especially databases. Rolling back code can leave it out of step with the database schema. You can sometimes ship a reverse migration alongside a migration, but that isn't always possible, and it doesn't answer what to do with the data.
Octopus's advice to customers is therefore to avoid talking about rollback at all: always roll forward. If version two has a bug, the fix is version three, not version one. This is what hotfix processes and fast feedback loops are for. Normally changes might pass through dev, staging, prod and approval gates. For a serious bug, the safest option may be to hotfix the current version and push it out as quickly as possible, with the build pipeline as the bottleneck and speed depending on the team's appetite for risk. If the failure was in the deployment mechanism itself, recovery can be faster still.
He said customers sometimes claim they roll back all the time. When asked what they do with a schema change, "they kind of stop" and realize it has been sheer luck that they never hit that case.
The host asked whether reverting is fine for stateless application logic. Erez said that ideally such a reversal happens through a feature flag, changing the code path instead of redeploying. Schema issues still need care, but the host noted, and Erez agreed, that the flagged code paths can be designed to stay consistent with whichever schema version is deployed.
Keeping feature flags from becoming a mess
The host raised the downside of heavy flag use, which they had seen at Uber: stale flags scattered across the codebase. Erez agreed emphatically. Octopus uses OpenFeature as its SDK and has built a wrapper so that each toggle in code records its owning team and an expiry date. Nothing breaks when the date passes, but the CI process notifies the team that the toggle may no longer be needed. Erez admitted that if he logged in, he would probably find notifications asking him to remove some of his own.
He compared this to weeding in the gardening metaphor for code: it needs constant attention. Some tools use observability data to show when a toggle was last evaluated, which is a useful signal. He also described a sequencing problem. A toggle lives in two places, the code and the flag platform. You should remove it from the code first, but that removal may take two weeks to reach production, and by then you've forgotten you still need to delete the configuration. Mechanisms that track the removal through to production and then report that the configuration is safe to delete help close that gap.
From shared test environments to ephemeral ones
Asked how development environments are evolving, Erez said there is no single pattern. Dev, test and prod is the most common setup, and even that simplifies things. In CD terms, "dev" is usually the first point of integration, where you check that the deployment works at all. Test is often kept close to production, sometimes with sanitized data, for QA and product teams.
He sees dev environments becoming less useful and ephemeral environments growing. An engineer working on a feature branch wants to confirm it works, let teammates see it, and then switch to something else, which is hard if it only runs on their laptop. An ephemeral environment spins up from the pre-merge branch with the dependencies it needs, deploys the app as if to a real environment, provides a URL others can use, and is torn down when the PR merges. This replaces arrangements like one test environment per tester, or a single shared test environment that testers must coordinate access to.
The host noted that talk about cloud development environments and preview environments seems to have faded, even though the technology exists. Erez said they are easy to discuss for simple cases but get tricky with multiple apps, stateful services and realistic data. The host suggested AI agents make them more valuable, since running the code and seeing it work, especially with a UI, is a strong form of validation. Erez agreed: whenever an agent validates its own work by testing and exploring, it is using an ephemeral environment, even if no human ever sees it.
Running Octopus Cloud
Erez is no longer on the team that runs Octopus's SaaS offering, but he gave some history. When Octopus first tried SaaS, around 2020 by his recollection, every customer got a dedicated VM with the self-hosted product installed. He estimated this cost about $100 per customer per month while customers paid something like $20. It was an experiment to test demand, and he credited the CEO for being willing to fund it. Moving from downloadable software to hosting was a big step for the company.
Demand turned out to be real, so a small group rebuilt the service on Kubernetes. Erez worked on getting Octopus running on Linux and in containers, while others built the dynamic worker infrastructure, partly to "stop losing money." The architecture uses what Octopus calls a "reef," in keeping with its nautical naming. A reef is a cell containing the resources for a set of customer instances, including a cluster and an Azure database, with some resources shared, and each customer instance runs as a pod. Today a dedicated team runs the service for several thousand customers and many thousands of deployments a month.
A current project aims to make deployments more resilient. Deployments run as imperative steps, and much of their state is held in memory, so upgrades require pausing tasks and restarting instances, which means some downtime. The goal is to get as close to zero as possible. Erez noted that cutting downtime from five minutes to one is "just work," while going from ten seconds to zero is a much bigger shift, and he isn't sure they will get there.
Why self-hosting isn't going away
Octopus rolls cloud updates out gradually, instance by instance, over a few days. On-prem adoption is much slower. When Erez looked at the numbers, it took about 200 days for 50 percent of on-prem customers to pick up a change and more than 400 days for 75 percent. Some customers run versions from five to seven years ago, so Octopus has to support upgrades across a range such as 2023.1 to 2026.4, with all the schema-migration baggage that brings.
Asked why Octopus doesn't just drop old versions, Erez said most of its customers are still on-prem: banks, financial institutions and governments that want full control of upgrades and downtime, whether on their own hardware or their own cloud accounts. He doesn't expect that to go away. Octopus has become more willing to deprecate old capabilities in recent years, partly thanks to adopting feature toggles, but he believes self-hosting will stay because for many organizations the cloud isn't viable or doesn't meet compliance requirements.
The host suggested that on-prem support may mean less competition for infrastructure vendors, and Erez agreed. The host also described long-unupgraded customers as potentially loyal ones. Erez related to that from his earlier job: if something critical like a deployment system works, customers leave it alone. That frustrates Octopus engineers who want to ship new features, but he praised the support team for helping customers on old versions. The host compared this to LLMs: when a new model version breaks a tuned workflow, some organizations may pay to pin a version or run it themselves.
Getting started, and what to read
For engineers wanting to move into progressive delivery, Erez's advice was to start with one feature toggle. Flipping something in production can feel scary compared with relying on tests well before code runs. But he said it becomes addictive, which is exactly why flag hygiene matters. He has shipped bugs behind toggles himself. The difference, he said, is that instead of panicking at 2 a.m. on call and wondering whether to build a new version or force a redeploy, you can turn the feature off, stop the bleeding, and then investigate calmly. Once you experience that, he said, you want to use toggles for everything.
His book recommendations were The Phoenix Project by Gene Kim, which their Skype manager gave everyone. Parts may be dated, he said, but its central idea of engineers owning the operational side of what they ship underpins everything discussed in the episode. He also recommended Radical Candor by Kim Scott for communicating directly while showing care, which he admits he needs as an engineer who can be blunt. For fun, he recommended anything by the Australian hard science fiction author Greg Egan, naming Diaspora and Schismatrix. He described Egan as a mathematician who builds entire stories from one altered physical premise, such as a speed of light that isn't absolute, and works out the consequences.
It's kind of funny how we talk about Kubernetes being cloud native. The reality is a lot of customers actually use Kubernetes for running on premise. I was talking to another one of our customers actually just the other day. They've got Kubernetes clusters running on research vessels.
Research vessels as in like…
As in boats.
On the ocean.
They've got Kubernetes clusters out in the open sea.
What is GitOps?
GitOps is potentially not necessary for all teams. Some of this absolutism that sometimes exists may not be necessary.
I don't hear too much chatter about rollbacks.
Rollbacks, this is always a spicy one. Customers can go, "Yeah, we roll back all the time." And then we ask them, "What do you do if you've got a schema change?" They kind of stop and realize that it's just sheer luck that they've never run into that. You want to avoid ever talking about rollback, it's always roll forward.
When it comes to CI/CD systems, what are you seeing changing there because of AI?
This is the elephant in the room. To be honest, it's…
CI/CD remains one of the hardest things to get right in software engineering. But why? Rob Erez is a CI/CD expert having worked in this field for more than a decade. In the early 2010s, we were teammates on the Skype for web team, and then Rob joined Octopus Deploy as one of the first engineers 10 years ago.
In today's episode, we cover progressive delivery in practice, canary deployments, blue/green, and why feature toggles are often still better. What is GitOps and why it's not about Git, and where that everything in Git mindset breaks down. Why you should prioritize rollbacks less and focus on roll forwards, and many more. If you want hard-earned lessons about CI/CD, progressive delivery, and what's coming as AI changes how much code we ship to production, then this episode is for you.
This episode is presented by Antithesis. Verify your system's correctness without human review or traditional integration tests and avoid bugs or outages.
Today's episode will be about CI/CD. CI/CD at scale is one of the hardest infrastructure problems to get right, and the teams who nail it know that the details very much matter. This is where I need to mention our season sponsor, WorkOS. WorkOS brings the same rigor as many of us use with CI/CD at scale to enterprise auth. SSO, SCIM, RBAC, production-ready, battle-tested, and built to handle real load and real compliance requirements. To add enterprise auth without the infrastructure project, visit workos.com.
Rob, it's awesome to have you here on the podcast.
Hello Guga, it's good to be here. Yeah, and loving loving An Jim.
Yeah, it's been like what? 11, 12 years since we worked together?
Yeah, yeah, I think 2014, 2015. I think I left UK. Yeah, that's it.
And Skype when there was still Skype on our team somehow in the heritage, the outlook.com plugin which had like 400 million users per month or something like that.
It was crazy the amount of money we were making.
So this was an interesting job. Deployments were very much a case of, you know, you ship once a week and you have to go to a CAB board, you know, a change advisory board and you have to get sign-off and approval. Yeah, and I always found that really weird, right? Like we're building this piece of software. It runs on the web. We can ship it whenever we want. It was running on Azure at the time and so, you know, we've got full access to push whenever we want and we make these changes through the week, but we'd kind of have to hold them back.
I guess to Abdella, our manager, both of our managers at the time, we kind of I guess worked around the system. When the code was ready, we'd build it and ship it through the week. I was really sort of impressed and proud at this process that the whole team had kind of put together, right? Where we'd, you know, commit the code, the test would run, several kind of layers of testing, we'd go to staging, etc. And then it would get shipped to production. So we're kind of I guess executing a form of, I guess, you know, continuous delivery at the time and we would then ship ourselves, you know, once a week.
I kind of always like to tell this story that at the time, you know, when we'd have a build ready to go, we do a form of canary deployments. And so this is where you kind of roll out to a small percentage of your customer base. And we always found that the customer base that would be our test subjects was New Zealand. So, New Zealand was always our canary.
Yep.
A bunch of reasons for that, you know, they're the first country to kind of reach this, you know, new date. So, they're always the first ones to kind of roll out into a time.
When it comes, like, you know, midnight passes, it's like 1:00 a.m. First country is New Zealand.
Bang, exactly. So, first country that's of, you know, significant size. They speak English, so if there's any bugs or issues or reports, it's kind of easy to understand. But, you know, to be honest, New Zealand is small enough that no one really cared if we shipped a bug and had to fix it quickly. So, sorry to all New Zealanders listening.
Yeah, I think that's kind of this good example of using a continuous delivery technique to, you know, ship the code faster than what we otherwise could have if we had these kind of big bang releases. And this whole process, I guess, opened my eyes to, you know, what progressive delivery, what good CI/CD could be. And yeah, I guess from there, I spent a few years there at Skype and eventually my wife and I decided it was sort of time to come home to Australia.
And then back in Australia, you went to start to work at Octopus Deploy?
Yeah, eventually I came back and actually worked at a place with a friend of mine just for a little while, just to kind of get back on the feet. And I remember they were using Octopus Deploy there. And so, Octopus Deploy, for those who don't know, is a deployment tool that's built and sort of developed originally in Brisbane, so there was a strong kind of Brisbane…
Attachment.
Attachment to it, yeah, that's right. So, when I found out that they were hiring, I thought, "Okay, why not? I'll give it a go. I like CI/CD. I like this space. I think there's, you know, a lot of interesting problems in this space." So, applied and joined in. At the time, I was employee number eight or nine or something like that. So, it was very much still a bit of a startup culture. Definitely not startup in the sense of, you know, Silicon Valley wild parties and, you know, ridiculous spending, but startup in the sense that everyone who I worked with was an engineer. Now, even Paul Stack, CEO, he's an engineer. This is kind of where it started from.
And so, we'd all be working on code together. You know, if someone had an idea, you'd have a bit of a chat about it and ship it. So, we were the marketing, we were support, we were kind of a bit of everything. And, yeah, obviously the company has grown a lot since then.
The company was focused, from the start, they were focused on deployments, right? Can we talk a little bit on… whenever I think about deployment, I always say CI/CD. Continuous integration, continuous delivery. Why was there a focus on deployments? And is that the same as continuous delivery?
Yeah, interesting. So, you're right. Like, quite often people talk about CI and CD as this kind of interchangeable…
They're either interchangeable or the word is CI/CD. It's like the name attached to itself. Like, it's hard for me to imagine a CD without a CI, continuous integration.
That's right. And I guess the way to look at it is, you know, you've got sort of multiple stages of maturity of software teams as they kind of move on their way from, you know, initially CI, which is continuous integration. This is the idea that…
Well, initially it's YOLO.
Initially it's YOLO. Yeah, that's right. Initially it's, to be fair…
Just deploy to prod or SSH into prod. What we used to.
We've all worked in places where we've done that and that's the starting point. So, you're right. YOLO is the first stage. The second stage is, you know, continuous integration. And so, this is this idea where you want to keep, you know, integrating, merging your code changes into a single branch. And you want to be continually running tests against it.
Now, continuous delivery is kind of the next stage where, you know, we talk about testing our code and there's, you know, unit tests, integration tests, etc. But what you also really need to test is your deployment process itself, right? So, continuous delivery is this idea, okay, you want to make sure that at any point in time when I click the button to deploy, I want it to go to production once we kind of get to this place.
The next step beyond that, which, you know, not all companies necessarily reach, is continuous deployment. Right? So, this is the idea that not only are your changes being merged together at the same time and ready to go, but they're also being shipped to production, essentially.
So, the stages we have is first you have YOLO, then continuous integration, then continuous…
Delivery.
Delivery, and continuous deployment.
That's right.
What is the difference between continuous delivery and continuous deployment?
The big difference, I guess, is the question of do your changes go out to production automated? Does it kind of flow through without any intervention, I guess?
And then for continuous delivery, they go out, but not necessarily to production, right?
That's right. So, that's why you'll have environments like, you know, dev environment or testing or staging or whatever. Now, it's possible that, you know, some parts of that process may also still be manual. You know, maybe you only update the test environment once a week so the testers can play around with it again. But the key principle is that you could kind of push it through sort of automatically the whole way through if you want.
And what teams would not want to do continuous deployment, right? Because it seems to me continuous delivery you kind of want to get to, because then you just get more and more feedback, right? But then it is kind of a good question, like should it go out immediately?
This is the question, you know, everyone always sort of asks, like it's almost ready to go out, why can't we just push it to production? As engineers you want to push it out as soon as it's ready, right? The reality is it doesn't really suit every company, right? So, you know, it may be the case that some companies really do still have, you know, review boards where you need to validate: is this good to go out? Particularly if you're in an industry that has a lot of regulation and compliance requirements. And they need to make sure that when it does go out to production, it's done at the right time with the right people available, etc.
It's not necessarily true to say that everyone should be going to continuous deployment, because that's, you know, sometimes just not viable for various reasons. But if you at least got to that point where you're sort of continually seeing your changes go through all the testing, you know, you're promoting it through the different environments, which is, you know, you're therefore testing the process itself. If you can only click that button to go to production once a week or whatever, okay, that's fine. You know, you've done a lot of that hard work. You've mitigated risk, which is what a lot of this process is about, right? Is feel the pain as soon as possible and de-risk anything that could go wrong right up until that last point.
So, I know you're deep into CI/CD, or continuous integration, continuous delivery, continuous deployment. You've been doing this for like what, 10 plus years now. But I was pretty surprised to see that when I checked Octopus Deploy, it says continuous deployment, continuous delivery, but it also says Kubernetes. How has Kubernetes kind of arrived in the topic of CI/CD and in general infrastructure? What happened there?
Yeah, Kubernetes is the platform of the moment. If we take a bit of a step back, Kubernetes came out of, you know, Google. I guess they originally had Borg. You know, they were using it to host and run their infrastructure. They ended up releasing Kubernetes partly, I'm not going to, you know, pretend I can read their minds and know exactly why, but partly as a way of helping to level the playing field between them and some of the other cloud vendors.
Yeah, so like before Kubernetes, AWS was a clear leader. And I talked with Kat Cosgrove who came to the podcast, who works on the Kubernetes team. And again, she speculated that by releasing Kubernetes, it was a lot easier to move workloads between AWS and Google Cloud. So, it kind of leveled the playing field and now there was a reason to…
That's right.
Like choosing Google Cloud was no longer as big of a risk, or choosing Azure was not as big a risk and so on.
Yeah, that's right. It made it simple to move between vendors, and so as a customer of one of these platforms, if you wanted to move to AWS and you're using containers, no problem. You just sort of put it in your place. And so Kubernetes came along at the time when there was a bunch of players in the field for container orchestration. So, you know, you had Nomad, even Docker Swarm…
And Docker Swarm.
A bunch of other options out there, because…
Also CoreOS. Kelsey Hightower was just on the podcast. They built Fleet, which was another container orchestration, and this was all around like 2012, 2013, 2014, and then Kubernetes came out and somehow started to win market share.
Yeah. Yeah, I mean I think some of the mechanics that it provided kind of really appealed to engineers and I guess DevOps teams out there. And eventually, I think particularly because it was so easy to use cross-platform, because some of the cloud vendors then did end up picking it up, it has ended up now being essentially the winner in the space. I know even back then, even non-container orchestration tools like Azure Service Fabric was another kind of attempt to handle the fact that, you know, everyone's building microservices and they want to host them in a single platform, and how do you orchestrate that and deal with dependencies, etc. But Kubernetes has become the clear winner.
And when you say winner, I understand that, for example, when you have a bunch of back-end servers on a service, like, you know, you have a website, there's a large back-end. Okay, I'll use Kubernetes for that. But you're talking about infrastructure, right? Or you're talking about even things like build servers.
That's right. So, it's kind of funny, you know, we talk about Kubernetes being, you know, cloud native. This is the term you always hear. It's cloud native.
Yeah, that's what they say, right?
That's what they say. And you know, you look at the vendors that picked it up. It's Azure and AWS, and they kind of made it available on their platforms. The reality is a lot of customers actually use Kubernetes for running on premise. So, you know, a non-insignificant number of our customers who are doing Kubernetes are running on potentially their own VMs on their own server farms. Or maybe they're running VMs in Azure or AWS, but they're maintaining Kubernetes itself. The idea being that they have a lot more control over exactly what's running. It's particularly common you'll find in things like the financial industry and things like that where, again, wanting to
fully sort of control the process and manage the whole sort of piece of infrastructure from end to end is kind of one of their goals, but they want to leverage the capabilities that Kubernetes provides by, you know, allowing the application team and the ops teams to just build and define kind of in that declarative fashion that Kubernetes provides exactly what runs and how does it run, etc.
So, they chose Kubernetes because this is the best tool they can manage their on-prem infrastructure and say like, "Okay, I have like these physical machines and I want this many virtual machines and I want to run a database on this many nodes and an internal web server or like whatever." So, it just won this area as well?
Yeah, yeah. I mean, so around the same time that Kubernetes came out, or actually before that, there were a lot of these other kind of declarative type tools, right? So, you have, you know, Terraform, which is a really popular one. You can define kind of exactly what infrastructure you want, and what you're doing is you're essentially defining the desired state. And then the tool kind of applies it, and you know, you've got Puppet, etc.
And so, Kubernetes has this similar concept, right? Where you define what you want your sort of infrastructure to look like and the internal Kubernetes controllers and operators will basically ensure that whatever you've asked for always applies. So, if you say you want, you know, three replicas of something, it will ensure that there's three replicas of something. And so, if one of those pods dies, for example, it will spin another one up. And so, it simplifies this process of being able to define declaratively kind of exactly what you as a sort of application team need to run your system.
It's fascinating because I always assumed that Kubernetes has won the cloud native space and hyperscalers. Can you tell me a bit more about how it's being used on-prem? Like some interesting stories. You must have seen some of it because you said that you're working with companies who are managing like large on-prem Kubernetes, or like interesting situations.
Yeah, this is one of the nice things about working at a company like Octopus, right? We talk to and deal with so many different customers. And yeah, everyone's doing things a little bit differently. They've got slightly different needs and requirements. And you kind of get exposed to a lot of different, you know, problems and patterns.
And it's easy sometimes to get lost in, you know, what people are talking about in conferences. And everyone's saying it's all about cloud and this is the, you know, best practice and you should be doing this. And the reality is, you know, everyone's kind of got their own little problems and they just want to solve them the way they kind of need to solve them. And so, some of our customers, in fact, a lot of our customers will run Kubernetes kind of, you know, quote unquote on premise. So, for example, I was actually talking with one
When you say on premise, can you just be a bit more clear? Is this a data center where they're like renting and co-locating? Is this actually like I have my own data center, or is this like I actually have my own machines in my closet?
Yeah, yes and yes, I guess. So,
What? You've been in a closet?
I mean,
I was trying to joke there. [laughter]
I'm sure there are still teams out there that are running, you know, the core accounting tools, etc. Yeah, under Steve's desk. But even when we talk about, you know, small computers, some of our customers have Kubernetes clusters basically in their point of sale systems. So, they have hundreds and hundreds of stores and they have little Kubernetes clusters essentially running them, and each one's independent.
And they, you know, run into their own problems with that, particularly at scale when you've got, you know, thousands and thousands of clusters, and, you know, these customers are following very GitOps practices, etc., where they're pulling the actual state from a Git repository. So, the Git repository itself becomes the bottleneck, or they start getting throttled. And so, they have to sort of resort to other mechanics to try to sort of mitigate and work around that.
I was talking to another one of our customers actually just the other day at KubeCon there, who they've got Kubernetes clusters running on research vessels. And those research vessels
Research as in like boats?
As in
As in ships, like on the ocean.
That's right. I'm not going to pretend to know exactly what they're doing on those ships. We didn't quite get into that detail. But they've got Kubernetes clusters out in the open sea, right? Which is apt, given Kubernetes' name. The problems they run into though are a little bit different, right?
So, for them, you know, those boats might be out at sea for, I don't know, weeks, months at a time, or whatever that might be. So, when you want to do a deployment, the ship's not available. So, when that ship comes back into port, it needs to get the update, right? So, we were talking through how you would achieve this, right? And how that process would work.
This is super interesting, and I love how you kind of get a peek into so many different types of teams through the fact that, you know, you're talking with them about how they do the deployments, but you probably see some other things that they're doing or things they're struggling with. What are some trends you're seeing across the industry in terms of this wide range of companies you work for, from startups to like finance companies to like these research vessels?
Yeah, I guess one of the big trends these days is a lot of focus on GitOps. So, GitOps is
What is GitOps?
What is GitOps? That's a good question, Gergely. Let's take a step back for a minute. So, you know, we talked about Kubernetes earlier. And we talked about the fact that it's kind of got this internal continuous reconciliation process where you say to the cluster, please spin up, you know, five pods, and it takes that desired state and it ensures it always sort of is true in the world. And so, there were a lot of products around that that were doing a similar thing, you know, Terraform does that for infrastructure, etc.
And a bunch of people started wondering, why can't we sort of take that process and pull it back further so that not only is Kubernetes just dealing with desired state, but we can pull it sort of directly out of Git? And so, you know, I as an engineer can make changes to that Git definition, that desired state, and I'll have some process that essentially pushes that to the cluster and ensures that it remains in line with what I'm expecting.
And so, the term GitOps was coined by Weaveworks in, I think it was 2017 or so. And as a general practice, it's just started picking up steam, particularly in tandem with Kubernetes, because at its core, Kubernetes is very declarative, right?
Later on, sort of in the early, you know, 2020s, it was kind of formalized a bit more, and there were sort of four key pillars of GitOps. The first being essentially declarative. You want your state to be declarative. So, this is the idea that you want to define what you want the state of your infrastructure to look like. This is to basically make things a lot, I guess, simpler to understand what the state of the world is going to be when a deployment takes place.
So, if you think about deployments that are a bit more imperative, that have sort of a process, the end result is sort of the result of multiple steps. But when you want to just update some infrastructure, that desired state kind of works really well, particularly in the Kubernetes space.
And then in GitOps, the desired state will be just describing like how many nodes I want, or like how many, I don't know, replicas do I want on a database, or how many web servers, or like how the load balancer is to be connected, that kind of stuff.
That's right. Yeah, so it's basically a way of being able to say I want my infrastructure to have whatever state it is. And then the GitOps agents, GitOps products, basically ensure that remains the case. So they'll keep applying it to Kubernetes. So you've kind of got this situation where Kubernetes keeps its internal status in sync with reality. And now you've got these GitOps tools that keep the declarative configuration state in sync with what Kubernetes is
So they will take whatever I put in Git, in whatever format I use, and they kind of translate it into something that makes sense for Kubernetes, and now Kubernetes can apply it.
Yeah, I mean ideally you want it as close as possible to what, I guess, Kubernetes is expecting, because
Or it allows you.
That's right. And so what you're describing there, I guess, is the continuous reconciliation. And so this is the idea that these GitOps apps will essentially, as we said, sort of take that state and apply it, and if there's any drift from the Kubernetes side, so for example someone, you know, runs kubectl delete pod or delete deployment or whatever the case might be, because your desired state is now stored in Git in this case, that will kind of self-repair.
The second pillar of GitOps is that that desired state you've sort of defined should be stored somewhere that's immutable and versioned. And so this is the idea that once I say that I want to have this state, I want to have sort of something I can point to or point at. And that might be a tag or a commit SHA or whatever. And I want to basically use that to define what that actual state should be. And I don't want that to be able to change, right? Because otherwise that kind of defeats half the point.
By having it versioned and immutable, it also makes things like auditing a lot simpler, right? You can see the transition of that desired state over time. What's interesting though is a lot of people will point to that and go, yes, versioned and immutable, I know what that is, that's Git.
I was about to say that, because Git definitely gives you versioning, or it gives you commit history. I'm not sure if it gives you versioning and immutable in the sense that, I mean, the past cannot be changed.
That's right.
Or actually can it? Because you can rewrite
Yeah, it's a
the history.
You're right. So, depending on how you sort of configure your GitOps agent, you know, you certainly can rewrite history. If you have it pointing at a tag, for example, you can change tags. And so, that's why there are, you know, best practices around that, I guess kind of, you know, wiggle the finger a bit if you're using tags to manage that sort of state.
But what's interesting though is really nothing in these pillars, and very quickly, the third one being pull versus push. And so, this is the idea that your GitOps agent will pull the state from GitHub, or Git, I should say, and put it into the cluster. And the fourth being continuous reconciliation. But nothing in any of these sort of pillars actually talks about Git. And I think that the naming of GitOps kind of gets people to already have this expectation that everything has to be in Git.
I mean, why would you not have that expectation? That's what I assumed.
That's right. I think the problem though is not everything should be in Git, right? So, you've got this constant kind of conversation within that community about, you know, where do you put secrets, for example? So, no one would say
Git, we would not. No, no, we know that, right? Do not put it in Git.
And so, that's the thing. So, you know, there have been all these solutions to try to put it in Git. So, there are things like Sealed Secrets, where you encrypt it and put it in Git.
Sounds like a terrible idea.
But I guess what this is really highlighting is the reality that some things don't need to be in Git, right? As long as you can have this sort of control over the versioning and immutability of it, then that's completely fine.
And then the trend around GitOps is what you're seeing that a lot more infra teams are moving from, okay, a few years ago they might have just made definitions for Kubernetes, and now they're moving over to GitOps, so saying, "Okay, we'd like to control infra in a tool, in a way that's described, that's in version control"? Is that the trend, or what is the trend around GitOps?
I guess it's more just the trend of the growth in general of GitOps in enterprises, right? So, not every company out there is using Kubernetes today. And as they sort of approach Kubernetes and they're looking at, "Well, how do I, you know, perform the deployments? How do I manage that process?" GitOps becomes the sort of de facto process.
And to some extent it is giving rise to this idea of using it to manage other things outside of Kubernetes. And there are a few examples of projects and experiments that will use things like Terraform, and there's a continuous reconciliation service that keeps your actual state outside in sync. At the moment, the focus is on, I guess, Kubernetes; that's the core place where it lives. And I guess it's more that the growth of Kubernetes itself means that GitOps is coming along for the ride.
And you mentioned enterprises, which means like these large companies with oftentimes thousands of people, or in regulated environments, so that's what I think of as enterprises. Are you also seeing smaller teams pick up things like GitOps? Is it like everywhere, or is it more that there are certain types of teams that seem to be just more interested in this?
That's a good question. So, I guess sometimes what we see is a lot of people go to conferences or they read blog posts and they hear that GitOps is what you should do. So, I guess what I want to point out here is GitOps is potentially not necessary for all locations, all environments, all teams, right? There's certainly a bunch of benefits to it, but the reality is there are some things you need to do outside of just GitOps. You might use GitOps principles in parts of your process, but some of this absolutism, I think, that sometimes exists may not be necessary.
So, there's often a bunch of other processes you do around your actual sort of, you know, quote unquote deployment. So, things like maybe you run smoke tests, or maybe you want to send a notification when it's complete, or maybe you want to do a database update or something like that. These kinds of steps don't really lend themselves very well to this declarative, everything-is-in-Git kind of process, right? And so, that's why you get things like Argo Workflows and Rollouts and things come out to try to kind of GitOpsify this process. And that works for some people.
But the reality, I guess, is that I think some people get really hung up on this idea that everything is Git, so therefore they've found the tool, and, you know, therefore everything is a nail. And I think that's not right.
And this is the thing, like talking with customers, when we go through this process of, you know, you can use GitOps in Octopus, and, you know, we've got a bunch of support for various mechanics that integrate well with Kubernetes and Argo, but there's a bunch of other sort of operations you do around that process that don't need it. And when you talk to them about it, you know, they realize that what they're trying to do is ultimately just ship software. So, again, that difference between what you hear when you talk at conferences and things, where, you know, everything is Git and everything must be, you know, in
in this particular format or whatever the case might be. The reality is for most customers, they're just trying to ship software, and they don't care what name you give it. If it's GitOps and it works end-to-end and solves everything, good. If they want to use GitOps as part of the process, but then have other mechanics that are more sort of imperative, then good. It's just sort of the reality of, you know, there's tens and tens of thousands of companies out there in the world that are doing software delivery. And not all of them are at conferences, and not all of them are at the, I guess.
As Rob says, most teams don't care whether you call it GitOps or anything else. They just want to ship software and know that it works.
Our presenting sponsor, Antithesis, does exactly that. It lets you ship knowing that it works. Antithesis goes beyond code review. It runs your whole system inside a hostile simulation. By doing so, it finds every bug before your users do. And because the simulation is fully deterministic, Antithesis doesn't only find bugs, it gives you a perfect reproduction of every issue. I know this sounds like science fiction, but it's actually hardcore engineering under the hood. Jane Street, fly.io, and the etcd community ship agent-written code with full confidence because they know it's been verified by Antithesis. To see more case studies in detail, head to antithesis.com/pragmatic. That's antithesis.com/pragmatic.
I also want to mention our season sponsor, turbopuffer. turbopuffer is exactly the thing that just works. A vector and full-text search engine built on object storage. Fast, cheap, and extremely scalable. No exotic architecture required.
Here's something I find interesting. The teams building the smartest AI products out there, Cursor, Notion, Cognition, Anthropic, they all run on turbopuffer. But, why? Let's think about it. An LLM without context, it can feel pretty dumb. I can still remember shortly after ChatGPT launched early 2023 how it felt both incredibly smart, but also frustratingly stupid. If you asked it a question that was outside its training data, it just made things up.
Fast forward to today, and the models and their tools for retrieving context are much better, but hallucinations still happen frequently. Here's a typical way to integrate turbopuffer with an LLM. Plug it behind your search tools, get fast responses and relevant context back. The neat thing is how it gives you the perfect blend of low-cost storage, fast retrieval, and a bunch of different search tools like vector, full-text, and filtering. You can afford to index billions of documents, and then you can query it with different tools to get the most relevant handful of documents in milliseconds.
This is how your LLM feels smart. It gets the right context really fast without blowing through a bunch of tokens. Of course, you can use turbopuffer to search anything, not just code. If you're building AI products, check out turbopuffer at turbopuffer.com/pragmatic. With this, let's get back to Rob and talk about progressive delivery.
Yeah. Another trend that we talked about just before is the rise of platform teams. Can you talk about what you're seeing?
So, platform teams are kind of, I guess in the past several years they've become this sort of new standard organizational structure to help teams manage their, I guess, deployment workflows, a bunch of infrastructure around it. And it's kind of come out of this evolution of DevOps, right? So, you know, we mentioned before, in the old days you'd write a bit of code and you'd throw it over the wall to the ops team. So, it was dev teams and ops teams and you
That was like in the 2010s, 2000s.
Back in the long, long ago. And then DevOps became, you know, the practice that everyone sort of realized that actually we want to have the engineering teams be involved in and have ownership of part of that operational process. Idea being, you know, you get faster feedback loops, you are able to kind of, if you feel the pain, you sort of fix... You know, it's that saying, you know, you fix it, you ship it. You know, we've all kind of heard that. And so, a lot of teams, you know, took that to heart.
That's good, great, good practice. But as things start to scale up, what you'd find is that there would end up being like a DevOps team again. And sometimes it's separate to another ops team. And so, there'd be this separation of development and DevOps. And it kind of goes against some of the principles of what DevOps was, you know, trying to destroy. But not only that, these teams then end up having lots of different ways of doing their deployment. So, you know, you've got a whole bunch of application teams and they've all got slightly different requirements and they're all building it from scratch. And so, you'd end up with these teams, whether you had the DevOps teams or it was still within the application teams, where there was just this large number of different ways of doing things, right? And that becomes difficult at scale. So, you know, you can't really move between teams.
And by scale you mean typically when there's a lot of teams, right? That's the easiest
Yeah, that's right. If you've got lots and lots of teams and each one is kind of owning that process end to end, you know, you sort of get this bifurcation of processes. And not only that, the application teams themselves start kind of getting this context overload, right? They now need to think about what's best practices of the different cloud tools.
Yeah, and devs rarely want to configure the deployment scripts and test them, and testing is hard. So, you can't stop with unit testable. So, it's now a different job. I remember when I was on earlier teams where, you know, typically on the mobile team you have a mobile team with five people and one of us had to kind of specialize in Jenkins configurations because Jenkins is oftentimes, or used to be, the mobile CI/CD. And it's kind of like half a person dedicated to that, and it was more like, you know, we had to draw a stick on who's going to do it, 'cause we want to build something.
You want to write code, right? You just want to focus on writing code, and so, if you're spending a bunch of your time sort of managing infrastructure and pipelines and things, you know, that's no fun for anyone. And so, platform teams have come about as a new way of solving that problem where it's different to, you know, this idea of a DevOps team or Ops team that kind of own the whole process. They more sort of define best practices and they provide ideally a self-service mechanism where application teams can essentially use often what's called an IDP, an internal development portal. And they'll be able to essentially self-service and, you know, maybe they want to spin up a new project and they're able to use a template that the platform team have generated. And so the platform team are able to sort of create these standards throughout the company.
And they can be responsible for, I guess, the definitions of those processes and the best practices and how to achieve that. But the ownership of the actual running operational sort of element is still within the teams, right? So they still get those benefits of, you know, DevOps being close to the real code and feeling the pain if there's a problem, etc., etc., etc. But they don't need to spend all that time becoming experts in, you know, all the different ways that you can deploy the software they've got. And so this has become really common now where, particularly as you sort of get to a larger size, platform teams are a great way of solving that problem.
Now, that's not to say that every company everywhere should have a platform team. You know, if you're a smaller company, sometimes you've just got the app team and they sort of are doing, you know, quote unquote DevOps. But this is certainly something that, as you sort of start seeing larger organizations with multiple teams and multiple projects, these platform teams are a way of basically bringing some sanity and control and focus, I guess, to the whole space.
One trend across the industry, of course, is AI. It's hard to see any teams where devs are not using AI agents specifically to code. Your product managers will be using these things, and of course we have a lot more code produced as a result. When it comes to CI/CD systems, what are you seeing changing there because of AI?
This is the elephant in the room, right? How is AI affecting sort of this? The reality is, I think, to be honest, it's still very early. I think what will happen is the impacts on CI/CD are really tightly coupled to how development teams end up using AI. So, there's going to be, I guess, a lagging process there. But, we're finding a lot of teams are starting to use AI in the development process. And so, we're starting this process of going out and looking and talking to customers and learning what's the way that they're handling AI in their teams and their application teams. And then how we can best leverage sort of the CI side to support that, but then in addition to that use AI within the pipeline itself. Again, in the right place.
So, one of the things we've been, I think, pretty keen on at Octopus is this idea that, you know, at KubeCon we're probably one of the few companies there that didn't have, you know, AI plastered all over it. Like we tried to be very... yeah, yeah. That's what gets the sales, right? That's what gets the sales here.
You stand out now these days.
That's right, by not having AI. I mean, we've got AI in Octopus, but what we've been trying to do is think about how do we actually use it in a way that's actually useful for our customers, right? For engineers, etc. And so, we've been slowly adding capabilities within Octopus to provide, you know, AI support, whether it's an MCP server, whether it's a recovery agent that can review logs and tasks and all that sort of thing. But, that's within the product itself.
Some of the bigger changes will depend on, like I said, how actual application teams use AI. What I think, yeah, what we're talking about, we'll find is there's going to be a lot more velocity. I think that's one of the big changes, right? Is there's just going to be a lot more code coming through. I think one of the questions is, okay, what does that mean for your pipeline?
One of the things you often talk about when there's a human element to the pipeline is speeding up the cycle to get that feedback quicker. You know, if you've got engineers sitting there waiting for their code to run tests so they can get back to it and fix it, the shorter and shorter you can make that feedback loop, the better it becomes because I don't need to context switch, etc. I think in a world where the majority of your code is being developed by AI, that becomes perhaps less important. You know, if you can kick out your build and test process and it takes 30 minutes versus 20 minutes, does it really matter if the engineers are already long gone, moved on to the next problem, and the actual AI agent itself can kind of babysit the process and review the problem that came up and issue a new fix?
I guess that there'll be a de-emphasis, I think, on some of the speed of the pipeline itself and more on decreasing risk, right? The risk that comes from having AI agents generate code. And so exactly what that process looks like, I guess, remains to be seen. I think what we'll see a lot more use of is things like progressive delivery. And I think particularly feature toggles are going to be a really common tool in the tool belt of application teams. Partly because it allows you to ship that code as fast as you can or as fast as you want, but manage the rollout of the actual feature set or changes sort of independent of the deployment. So it decouples your deployment from your release. And so in a world where, you know, we've got a lot more AI agents generating code and being involved in perhaps part of the build process, those agents themselves being able to use toggles to react to it quickly, I think, then become a lot more important than perhaps what we see today.
Can we talk about progressive delivery, what it is, and what are the most common ways to, you know, de-risk getting your code or your software out there?
Progressive delivery is the next evolution beyond continuous delivery. So, you know, with continuous delivery, it's this idea that, you know, I've made a change to the system and I want to ship it to dev or staging or typically, you know, if it gets to production, sort of in one hit, right? With progressive delivery, what you're trying to do is basically release those changes in a little bit more of a controlled way, typically through things like a canary deployment. So, this is where you might deploy to some subset of your instances that are out there.
So, what is a canary?
What is a canary? Canary deployment is... This is New Zealand, basically. New Zealand's our canary. So, this is, as we said before, this idea where you select some subset of your customer base or whatever that might be. And you would typically route traffic to a new instance. So, you know, you've got version one running and you want to release version two. You essentially ship version two side by side, and you might use, you know, the most common one would be some sort of network traffic manager to route some percentage of your traffic to that new instance. And you gradually roll that up. Typically, you know, as you sort of do this process properly, you should have fairly mature observability mechanisms in place to see that, you know, you can roll up or roll down.
And I guess this whole thing comes from a canary in a coal mine, right?
That's right. Yeah, yeah, so the idea being that, you know, in the old days when you'd be in a coal mine digging away and you would release, you know, all sorts of toxic fumes and things like that. Canaries were a lot more sensitive to it. So, they'd have a little canary in a cage. And if that canary sort of died, I guess, got knocked out.
I think the canaries, as I understand, they were like tripping. And then
Okay, oh, that sounds better.
But when it stopped tripping, well, it also died.
Oh, okay. So, all right. Same ending, but just, you know, a nicer way to go out.
Then you just get out.
Yeah, so it's this idea that you get that advanced warning, I guess, that, you know, rather than you getting knocked out by the toxic gases, etc., you know, you can get out of there sooner. So, it's that same principle, I guess, brought to the software.
There's various other mechanisms like blue/green deployments. So, you've got your first version that's still receiving traffic, and your second version is up and running, and you can now do some tests against it, validate it. Maybe you've got, you know, the IP details to access it directly. You can basically validate it's working. Sometimes there may be a way of avoiding cold starts and things because that process may need to, you know, initialize a bunch of stuff. But then when you've sort of done that validation and you're ready, you can essentially swap traffic around. So, all the new traffic goes to the other. In some ways it's like doing a canary, but straight to 100 percent, but you're doing a bunch of validation sort of on the side before it actually reaches customers.
In my view, probably the more useful progressive delivery strategy is feature toggles.
So, this is the idea that you've got some sort of—
Feature flags as well, right?
Feature flags, feature toggles. Yeah, that's right.
Yeah, often used interchangeably. So, this is the idea that you've got, you know, some sort of variable in your system, and it's linked to typically some sort of external service, and through the state of that particular variable being true or false, on or off, you're going to essentially have different code paths take effect. And there's a bunch of benefits that feature toggles have over, I'd say, canary releases, particularly for application delivery, where your unit of change with a feature toggle is very granular. It can be, you know, single lines of code. And so, everything else remains the same, and all you're doing is tweaking that single line of code.
With a canary or any sort of version delivery deployment mechanism, your unit of change is the entire app. So, if you've had, you know, 20 commits since the last release went out, then you're essentially testing all 20 things in that one hit. Your ability to target the actual customers, it's a lot more precise when you're using feature toggles. So, you can, you know, use all sorts of complex rules, and say that, I don't know, everyone from Germany who has this particular product in their basket has this kind of experience. And that's, you know, really hard to do via network traffic rules, right?
The other is your ability then to actually roll back. So to roll back from a canary, hopefully you're still in the process where you're going through that canary process and you can roll it back. That could take, you know, minutes. Maybe you have to redeploy the whole old version. That could be minutes or more. With a feature toggle, you know, you can do that in seconds. That's pressing a button and it happens immediately.
Not only that, but you've got more control, I guess, on when you do that. So, with a deployment that you're doing via a standard version release, you're tied to when that deployment takes place. Cuz when it takes place, that's when essentially your new feature is available. And as an application team, that means you need to know about exactly when it's taking place and make sure you're watching the logs at that point, and maybe you and 10 other teams who are shipping things at the same time are all doing the same thing. Whereas with feature flags, you've basically got control over when that takes place. So, you might ship the actual, you know, the assemblies and all that sort of thing on the Monday, but you release your feature on Tuesday when you come in and you've got the logs ready and you've kind of reviewed what the next steps are. So, this really makes things a lot easier to decouple releasing a feature from deploying software.
You know, version deployments through canary, etc., they're really useful, particularly if you're doing like infrastructure type changes where there is no kind of application toggle that's relevant there, but you want to validate some changes to your infrastructure or your process. Or potentially, you know, things that will involve schema changes. And schema changes are the big— the big schema changes, right? That's a schema change. This is the big problem in any, like, to be fair, in any progressive delivery.
And this is why, you know, the question always is, are you ready for delivery? To do schema changes. But I guess this is the point that application teams kind of need to be really mature, and I don't mean mature in terms of, you know, not telling silly jokes, but mature in terms of understanding all the problems that are in place with this and knowing how to release these sort of changes in a gradual, controlled fashion and do it over multiple stages.
That, you know, ironically is actually quite hard for us at Octopus because our software is both SaaS hosted, so we have a SaaS offering that customers can use, and we have an on-premise version. And because we have both sides, we kind of have the best and worst of both worlds, you know. In the cloud system, if you've got a SaaS product, you have complete control over what versions go where. So, if you want to do expand and contract, you can stage the whole process. You know that it's all been updated before you move to the next stage. On the other hand, for a self-hosted application where they go in and they install it on their own infrastructure somewhere, you don't know what version they're running and what they're coming from. So, they might go from version one straight to version six. And so, you're not really forcing them to go through that expand and contract phase.
On the other hand, you know, they've got a lot more control over when they upgrade. And so, you can be a little bit more deliberate about, you know, making sure that they do backups before they change and, you know, maybe they can manage that migration and accept a little bit more downtime during migrations and updates and things like that than would actually be, you know, acceptable in a SaaS product.
So, one thing about, you know, we talked about progressive delivery, and you're kind of doing this to avoid surprises. You know, if a regression goes out, a new bug, or something doesn't work, you kind of want to catch it early. Hopefully, only a few customers have experienced it. Or even if it's not 100%, you have a way to go back: all you do is, if it's a feature flag, you hide it. If it's a canary deployment, you go back to the other one.
But, there's also this thing where, like, when things do go wrong, at some point you want to do a rollback. Can we talk about how you have seen rollbacks done well, and what does it take to actually have a real rollback strategy? A bunch of people talk about CI/CD, some people talk about feature flags. I don't hear too much chatter about rollbacks.
Yeah, rollbacks. This is always a spicy one. We get a lot of customers who say, "Why don't you have a rollback button? I want to roll things back. Why can't we roll things back?"
As in, in the deployment software like Octopus or anything else, they're like, "Okay, if it can deploy, I want to do checkpoints." And just do a rollback.
How hard could it be? Just do what you— How hard could it be? Tell me.
How hard could it be? Well, this is the problem, right? So, in a completely stateless system that's, you know, pretty straightforward. If you've got a completely stateless system, and, you know, this is something that GitOps is really good at, where you'll have that definition stored somewhere in a repo. If it's completely stateless, you can do a Git revert and push it, and it'll go back. The reality is for most systems out there, you've probably got some state.
State being—
State being databases. It could be, you know, any sort of information that you can't necessarily just undo, I guess, because if you roll it back, now you've got your code talking with a schema of the database that's not in sync. If you've got a schema migration, let's say in a normal deployment, you can provide alongside that a secondary sort of anti-migration that undoes the change, but again, that's not always possible. What are you going to do with that data store?
We've gotten pretty far in basically trying to advise customers that you want to avoid ever talking about rollback. It's always roll forward. So, if there's a bug—
Okay.
Roll forward. Get a change in. Yeah, get your change in as soon as possible. This is where fast feedback loops are important, right? You know, this is what the hotfix process is all for, right? So, we all know that in a standard process, you want to go dev, staging, prod, and maybe you've got approval processes and it slows down, etc. But if you've got a significant bug that you need to quote unquote roll back, sometimes the safest thing to do is actually make a hotfix to that version and push it out as quick as possible. And your bottleneck might be the build pipeline or whatever, but depending on your appetite for risk there, you can resolve that a lot quicker. Now, obviously, if the failure itself is just from some mechanism in the deployment process itself or somewhere further down that chain, then your time to recover is going to be a lot quicker.
But it's this idea that, you know, if I've got a failure in version two, my rollback isn't to go to version one, it's to go to version three and make sure I've got that fix in version three. It's the sort of thing that, you know, when we talk to customers, some of them go, "Yeah, yeah, we roll back, you know, we roll back all the time if there's a problem." And then when you ask them, "What do you do if you've got a schema change?" they kind of stop and realize that it's just sheer luck that they've never run into that, right?
Is it fair to say that you want to roll forward if it involves business logic or something that is not stateless? Cuz if it is stateless or if it's application logic, you know, you have code that says, "If this, else then," and you realize there's a bug there, you can just revert it as long as it doesn't, you know, touch the schema or the data.
Yeah, I mean, in an ideal world, your reverting is through a feature flag, right? That you click and you're essentially reverting by changing the code path. And this is why I always say feature flags are a nice tool to use for doing this progressive delivery, because, you know, just as easy as it is to roll out that feature, you can typically roll it back. Now, you're still going to have some of those problems with schema issues, etc. If, you know, you're making a change and you've got parts of your code path that expect one and not the other, you're going to need to account for that.
But you can even account for that inside the feature flag.
That's right. Yeah, so that's the way you ideally manage that, so that with regards to which path you go down the feature flag, it's self-consistent with whatever version of the actual database schema that's out there.
So, I guess the more feature flags you use, the fewer surprises you might have. But, it's a bit of extra work both to build and also to remove.
Yeah. Yeah.
You get stuck with stale feature flags all across your code base when you start using it a lot. I saw this at Uber.
Yes. Yes. A hundred times yes. So, when you're adding a feature toggle to your app itself... So we at Octopus, we obviously use feature toggles in our code quite a lot. And we use OpenFeature as the framework, the SDK, to interact with it. But, we essentially have built a wrapper around it where, for the toggle itself within the code, we provide some details about which team owns it. And that team sets an expiry on it. Now, the expiry itself, when that time passes, nothing bad will happen. But, through parts of the CI process, if that time has passed, we can send a notification to that team and say, "Hey, it looks like this toggle is no longer used." So, the specific mechanics don't matter as much, but it's more a matter of making sure that, you know, if you're adding feature toggles, it's really easy to forget about it, because you start rolling it out and you can kind of forget about it. And, you know, you want to keep it in there just in case for a while in case you need to roll it back. And having the ability to understand how long a toggle has been there is kind of a key part of helping to maintain that hygiene.
Now, the reality is even at Octopus, we've got a bunch in— I know I've got a bunch in there that I'm sure if I was to log in, I'd probably get a bunch of notifications to remove. You know, when we use that gardening metaphor in code, right? This is one of those sort of operations. This is weeding, right? You need to just keep on top of it. There are some mechanisms around, even in lieu of the AI side, which will— you know, ideally if you're using feature toggles, you've probably got a bunch of observability and metrics and logging around it. And there are some tools out there that will allow you to keep track of when the last time a toggle was evaluated, and that kind of gives you that signal.
Similarly, you know, you might remove it from the code, because typically when you want to remove a feature toggle, you want to remove it from the code first before you touch your actual toggle system. And so, having a mechanism so that once you remove it from the code— you know, it might take 2 weeks before it makes all the way out into production. So, you don't want to delete it before then. By that time you've kind of forgotten about the fact you removed it.
And so, having mechanisms that will keep track of that change, I guess, going through the system, and when it reaches the environment where, you know, production, where it's actually being used, can kind of say, "Okay, that code's gone out. That's, you know, removed the toggle. It's fine and safe to actually remove the configuration." Because you've got that feature toggle information in two places, right? You've got it in the code and you've got it in your platform.
Can we talk about how development environments evolve? We talked about CI/CD, but I'm interested more in, you know, you go from like you have one environment, later you might have staging or something, and what evolution have you seen across all the teams that you work with, all these hundreds or thousands of teams?
Yeah, I'm not sure if there is one particular pattern there. I mean, I think, you know, most common is, you know, dev, test, prod.
So, these three different environments.
Yeah, and I mean, even that I think is probably a gross simplification of all the different kind of mechanisms there.
And dev meaning my local machine?
Dev in the case of CD is often the first point of integration. So, it's kind of— test, often customers will keep test reasonably in sync with, let's say, production or some sort of sanitized data source. So, that way, whether it's the QA testers or the product team or whatever, they can go and review the code. Dev is almost like the first point of integration: is the deployment process just at its core actually working, or is anything fundamentally broken at all? I think more and more now we're finding that dev is less useful in that respect. And what we're seeing is more the growth of things like ephemeral environments. And so this is the idea that, you know, I'm an engineer, I'm running some sort of feature on a feature branch, and I want to evaluate that it's actually doing what we're expecting it to do, but not only that, I want the rest of my team to be able to see it working, and, you know, if I've got it running on my machine, it's not exactly easy to, yeah, give other people access, I guess.
And then I may want to, you know, completely context change, move on to something completely different. So ephemeral environments is this idea: from my branch, pre-merge, I want to spin up a whole environment essentially from scratch, ideally with whatever dependencies are required to run this particular component that I've been building. And then I want basically to deploy my app into that as if it was a normal full-fledged environment. Once that's available, I want to have access to, you know,
If it's a web app, maybe it gives me the URL and I can poke around it and hand it around and other people can kind of evaluate. And then the moment I kind of merge that PR, tear it down again. You know, it's quite common to have multiple test environments because, you know, I've got a lot of stuff going through my pipeline and I've got three testers, so let's have three environments, so they can all sort of have one at once. Or often you'll see a single test environment and a bunch of testers, and they will kind of need to collaborate to see who's got access to the system at the moment, etc. etc.
Whereas with ephemeral environments, it doesn't roll off the tongue. With ephemeral environments, you can essentially have a full-fledged deployment per feature. And so again, that's about speeding up that feedback process, right? Again, these are all about speeding up that feedback process to catch those failures or issues or bugs or whatever sooner.
There was a time a few years ago where cloud development environments were really talked about a lot, which was the idea is as a developer you have an environment spun up in the cloud, let's say your Visual Studio Code connects to it, or maybe you just log in online, and it spins up all the dependencies, oftentimes done with containers, which reminds me of this as well. And there's also like preview environments, but somehow it feels that both that discussion and this one kind of died down. Maybe it's AI, maybe it's something else, but I mean, the technology is there, right? We have containers, you can package things together, I'm sure it depends, but it's all doable.
Yeah, it does get tricky. This is again one of those sort of things that's really easy to talk about for simple cases. It can get tricky when, you know, what if I've got more than just a single app in my kind of quote-unquote environment? And how do I make sure it's got all the data I need to validate? So, it can get tricky. So,
Or if you have a bunch of services that have state.
That's right, exactly. So, there are sort of complications that it does bring, but I guess the benefits that you get as an application team, particularly, you know, an application team where you've still got engineers writing code, is sort of speeding up that feedback process, I guess.
Well, now with AI agents everywhere, that's even better, 'cause in the sense that one of the best ways to validate, you know, we have code reviews and an AI and you look at the code, but isn't it better to just confirm that this thing works, especially when it has a UI?
That's right. I think even in that world where you've got AI agents kind of building the code and validating the code, any sort of scenario where you want that AI agent to kind of validate what it's done, you're essentially talking about ephemeral environments, even if it's not exposed to people, because it's doing its own testing and poking around in whatever shape or form that it's doing. That still is, I guess, one of these kind of environments, right? So, it's ephemeral, it spins up, you've got some sort of provisioning process, and then, ideally, once the job's done, you kind of tear it down.
I'm interested in learning more about the reality of operating a large infrastructure platform, and, you know, one big one you're working on is actually Octopus Deploy's SaaS offering. How does that look like, and what are the challenges of, you know, running something where you're running all of these deploy processes and all these CD? You probably have a bunch of different things. What is it like?
So, at the moment, I'm not on the team that sort of built that, but I can give some of the context, I guess, from history. Originally, when we first sort of decided to provide an Octopus SaaS offering, I don't know, I think it was 2020 or something like that, it was all VMs, so every customer would basically get a VM spun up, and we would
Virtual machine.
And Octopus, you know, self-installed, would basically get installed onto that VM, and they'd get a whole VM for running workloads on, etc. And that was very much not cost-effective. It was costing us something like 100 bucks per customer per month, and they were paying, I don't know, $20 a month or whatever it was. But this whole process was more an experiment to see, was there a demand? And to his credit, Paul was happy to, you know, pass out the credit card to kind of go through this process to see, is this actually the direction we want to go? Is this something that's going to turn into a viable sort of direction for the company? 'Cause it's a big step, right, going from building software that you can kind of hand out, and people can download and manage themselves, to
So, it was pretty much self-hosted, or, like, run on your own infrastructure, right?
Exactly. Yeah, that's right. And so, the demand was there. So, not long after that sort of first experiment, we basically started from scratch again, and I worked with a couple of the other engineers back then to start building it on Kubernetes. And so Octopus itself, in that space, we have what we call kind of a reef. What you find is everything in Octopus, we've always got Octopus or nautical kind of names around it. So, a reef is basically this cell-based architecture where it contains all the resources that are needed for that particular customer's instance. Well, some of it's shared, but it's kind of broken down into individual cells. And so, a reef will contain, you know, the cluster and Azure database, etc. And each customer instance is running now in a pod in that cluster.
And so, as part of that project, I think I was working on converting it so it could run on Linux and inside containers, and someone else was building the dynamic worker infrastructure. So, there were a couple of us that kind of just got in and, yeah, really just got it up and running so that way we could kind of start moving forward and, I guess, stop losing money. Fast forward to today, now there's an entire team that kind of backs that, and we've got, you know, several thousand customers on it, and we run, you know, many, many thousands of deployments every month.
And so, now what we're trying to do is there's a project at the moment to basically make the Octopus deployment process itself more resilient. So, what that means is, at the moment, when a deployment kicks off, a bunch of the process, so it's kind of an imperative set of steps, a bunch of that is stored in memory. Which means that whenever we want to do an upgrade, we need to essentially stop running tasks for some period of time so we can kill their instance and spin another one back up.
Octopus itself at the moment doesn't sort of have zero downtime between upgrades. So, there's a bit of downtime between that, but we kind of want to reduce that and get that as close to zero as possible, with the realization that, you know, going from downtime of 5 minutes to 1, that's just work, right? You can move things around, you can maybe change the architecture. Going from 10 seconds to zero is a much bigger shift. I'm not sure if we'll ever get there, but yeah, there's definitely this big effort at the moment to make the whole process a lot more resilient, to basically improve and reduce the amount of downtime that takes place, so we can kind of perform upgrades quicker, etc.
One interesting thing you do is you have a SaaS, but you also have an on-prem offering. What are interesting engineering challenges that come from that? A lot of companies have decided to just, honestly, move to SaaS, because now they control everything centrally. I think Jira did this, or maybe they're doing it, which is a well-known one, but clearly it's just a lot more work and a lot more headache to have both.
Yeah, and we touched on one of the big problems here a little earlier, which is that when we want to push out any updates, you know, to cloud, because we control the whole process, we can push it out. And so, we have sort of a gradual rollout process there. Because each customer is on their own instance, we can sort of deploy each one individually. And that may take, I don't know, a few days to roll out a change.
On-prem, though, is kind of another matter. So, actually, I was digging into some of the stats around this a little while ago and found it took about 200 days for an average 50%
50% of our customers on-prem to get it. Let's say I ship a new change today, it takes about 200 days for, on average, 50%.
It's half a year.
But then there's kind of like almost an exponential decay there, where it takes 400 and something days for 75% to get it. So, there's kind of this curve where, I mean, we've got customers that are still running, you know, versions of Octopus from 5, 6, 7 years ago. And so, whenever we ship a new change, we need to basically make sure Octopus will work from version, you know, 2023.1 to 2026.4. And so, there's a bunch more baggage, I guess, that we have in terms of
Schema upgrades, and making sure that whole process actually is achievable.
But why do you do it? A lot of startups will be like, screw it, let's not support old versions. This even happens on mobile. What's the benefit? It feels like you're kind of swimming against the current with this one.
The majority of our customers are still on-prem, and so, you know, you're talking about banks, financial institutions, governments, things like that, where they want full control over the system. They want to run it on their own hardware. Now, they may use their own cloud or whatever to run it, but they want to manage the whole process and be in control of, let's say, upgrades or downtime or things like that. So, it's certainly not uncommon, and I don't think that's going away anytime soon.
As for the upgrade support, we're kind of going through this process, actually. In the past couple years, we've been getting a lot more, I guess, confident with deprecating features and things like that, and just kind of cutting loose old capabilities. And part of that has come from, you know, fully embracing feature toggles as part of that process. I think we're getting a little bit braver in terms of, you know, removing capabilities that perhaps older customers may miss, but I don't think that in the long term self-hosted will kind of go away. This is one of those sort of things again where I think it's really common to hear, you know, everything's in the cloud, we're all in the cloud. And the reality is there's a lot of companies out there where for them it just doesn't make sense, or it's not viable, or it doesn't meet compliance requirements, or whatever the case may be.
Also, it's kind of a reminder, I think, that you actually might have a lot less competition if you build infrastructure software that also runs on-prem, because it sounds like there's demand where companies are like, we want to give you money in order for us to run on-prem. And I'm sure some of them would do SaaS if there's no other alternative, but SaaS is easier to build anyway, so there'll be more competition. So, if you're an entrepreneur, or if you're a software engineer thinking to do a business or start a business, it might give you an edge.
Yeah, that's right.
It sounds like a lot of your customers, you know, the ones who have not upgraded your software for, let's say, 5 years, you'd be like, "Oh my gosh, what are they doing?" But they might just be happy with it. And if they keep paying you as a business, those are some of your most loyal customers. You see what I mean?
That's right. And this is the thing. I mean, I remember when I worked in the previous job that used Octopus, or any of us who have any other sort of, you know, software that you've got running, potentially running locally: if it just works, why touch it, I guess? And so, it's kind of the bane of our existence, 'cause it annoys us. We want to ship the features and give them all these great new things. But on the flip side, you know, particularly with something as critical as the deployment system, a lot of customers, once they've got it running, they kind of step away and go, "Okay, let's just let it be."
And it keeps happening with AI as well, in the sense that, for example, I just read that Cursor's latest coding model is updated, I think, every 5 hours, which is amazing. It just keeps getting better. However, you know, there are customers who, once you have an LLM and it works for you, you've kind of tuned it, you have the instructions, great. But oftentimes what happens is a new version of a model comes out, or a major version, and it stops working. And I assume that there will be more teams, companies, businesses who are like, "Look, it would be worth money for me to kind of pin this thing, or to run it on my own infra and just have it stay as is, and then I will decide when I want to change it. If it ain't broken, don't fix it."
That's right. And I think, to Octopus's credit, we have a really good history of sort of helping customers even when they're kind of on those older versions, sometimes to the extent of wanting to say to the support team, they're on an old instance, tell them to get the fixed upgrade. But the support team are, you know, I think second to none in terms of their willingness to help. And as you said, if they're willing to pay us, who am I to say no?
Yeah, I mean, it's a business strategy, but I think it's just a nice reminder that there is not just one size. And even though I think SaaS is eating the world, and we're hearing it and we're seeing it, it's nice to see that it's not just that. As closing, if I'm a software engineer and I would like to move beyond continuous delivery, continuous deployment and go into progressive delivery, what pointers can you give me?
Yeah, I guess just start with something, right? So, start with adding one feature toggle. It may be scary at first to kind of go, oh, it's in production, if I toggle this, I'm going to break something in production. You know, it's nice and comfortable to know that you're kind of well to the left of the running systems, and if you ship code, everything will be caught by the tests. But, you know, if I toggle it, what will happen? It's kind of like a drug, right? Once you start doing it, you don't want to stop, and that's why we've got this hygiene problem for things like feature toggles, right? It's really easy to add them and actually end up with the opposite problem of how do you kind of control yourself? How do you stop? So, I'd say just kind of start doing it. Add one and keep an eye on it as you roll it out and look at the results from it.
And the reality is, you know, I've shipped features behind feature toggles where I've shipped a bug, right? And it's one thing to ship something and turn on a feature and go, okay, cool, customers have it. It's a very different thing when you ship something and there's a problem and you can reach immediately for the toggle and switch it back off. You know, the amount of times in the past you have this kind of panic of, oh no, I've shipped something, I don't know what's going wrong. And particularly when you're in that state, you know, maybe you've got called up at 2:00 a.m. because you're on call, and you don't know what the next step is, and you've kind of got a panic mind: should I build a new thing, or do I somehow force a redeployment?
So, having the capability of being able to sort of flick that switch just allows you then calm right down and go, "Okay, I've stemmed the bleeding. Now I'll come back and reanalyze it and understand what's wrong." And so, having that capability, once you sort of experience that and realize the value of that, not just rolling things out, but sort of I guess rolling that individual feature back off, yeah, you'll want to use it for everything.
What's one or two books you would recommend and why?
I'll give two kind of I guess technical ones and more of a fun. Phoenix Project is still for me a good one. This is one
By Gene Kim, yeah.
Yeah, and I can see, you know, you kind of remember that we got that in Skype. This is one that Abdallah kind of gave to everyone and
Yeah, our manager gave it to everyone.
Yeah, and you know, parts of it may be a little bit outdated and some of the practices have changed a little bit, but at its core this idea of as an engineer being involved in that whole sort of operations side of what you're shipping, and the value that gives to not just the company, but to you, is amazing. So, I think that book is kind of a cool one. Yeah, it's one of those cool foundation ones that sets the context for everything we talked about today.
The other one, from a more I guess organizational and communication side of things, Radical Candor by Kim Scott. How do you communicate more efficiently and with more compassion with your peers and other people around you. It's, you know, really common. I'm an engineer, so I know some of my bits really, you kind of look back on what you said and you feel like, "Okay, maybe I can be a little bit blunt." Whereas Radical Candor teaches us to think about, you know, you want to have those communications that are both sharing that you're caring and empathetic, but also direct. And you know, the benefits of that and kind of the inverse of that where, you know, you perhaps, like I said, you're very blunt. You're sort of being honest about it, but you're missing that empathy. So, I found that book really useful and interesting as, I guess, not even just as an engineer, but as a person working with other people.
From the more fun side, basically anything by Greg Egan. He's an Australian sci-fi author. He writes some pretty crazy and mind-bending hard sci-fi. So, if you're really into that, I'd say read like Diaspora or Schismatrix. They're the sort of books that actually took a second read to get through, and he's a mathematician as well, so yeah, he's got a whole bunch of background mathematics on why a certain part of his story goes the way that he... He wrote an entire story on the premise of what if the speed of light wasn't absolute or something like this. This one premise, and it kind of breaks out into, and then this is what happens to energy, and therefore molecules work like this, and da da da da, and as a, you know, I'm a tech nerd. That sort of science stuff, you know, really appeals.
So, the same, when sci-fi has some science involved, I find it way more fun.
Mhm.
Rob, thanks very much.
Thank you, Gareth. Been great.
What an interesting conversation. I hope you enjoyed having someone like Rob who has been building and thinking about CI/CD at scale for a decade. It was such a fun blast from the past story as he talked about how at Skype our team basically did continuous delivery years before most of the industry caught up, and how we did it by quietly shipping new builds to New Zealand every week, using this as our canary country. It's a reminder that a lot of modern software practices were already being run in the wild by devs who just wanted to ship software faster than our change advisory board would allow us to do so.
One other thing I took a note of is Rob's take on rollbacks. Lots of engineering teams talk about rollbacks as if they're a safety net. But the moment you have a database schema change in the mix, what safety net? Rob's advice is to roll forward, not back, and use feature toggles as a way to turn features off or on. This is also a reminder that investing in feature flags is usually really helpful. But if you have feature flags, be sure to clean them up after you've rolled them out, otherwise they become a big mess.
Finally, a part where I learned something new was on GitOps. I've always assumed that GitOps was about, well, Git. But as Rob pointed out, none of the four actual pillars of GitOps require Git at all. The four pillars are, number one, declarative. Number two, versioned and immutable. Number three, pulled, not pushed. Number four, continuously reconciled. The name GitOps has caused the whole industry to get a bit dogmatic about putting everything into a Git repo, even things like secrets, which absolutely should not be there. Rob's take is that most teams just want to ship software. If GitOps helps with that part, great. But if a more practical process works better, just use that.
Do check out the show notes below for related The Pragmatic Engineer deep dives on back-end technologies and other related topics. If you've enjoyed this podcast, please do subscribe on your favorite podcast platform and on YouTube. A special thank you if you also leave a rating on the show. Thanks, and see you in the next one.
Article published
