The As-If Rule, From Ractors to ZJIT to AI Agents: Aaron Patterson's Rails World 2026 Closing Keynote
Ruby on RailsAaron Patterson (Tenderlove) closed Rails World 2026 with a talk about performance, organized around one question: when an optimization changes what a program does, who can tell? His answer, drawn from compiler design, is the "as-if rule." An optimization may change anything as long as the behavior an observer can see stays the same. He applied that idea to memoization under Ractors, to stack frames that Ruby 4.0 quietly removes, and to allocations that ZJIT eliminates. Near the end he added a second story about AI agents chaining a RubyGems.org bug into an attack. His conclusion was that compilers have made us a promise about observable behavior, and AI has not.
Opening Jokes and a Rails Core Announcement
Patterson opened with a run of jokes about AI-generated software. He claimed to have "vibe coded" his own Linux distribution, "Aaron XP," which installs in 100 milliseconds because he reinstalls it constantly. He said he would start selling "insurance on technical debt." He called it ironic to joke about the financial crisis and then encourage people to generate lots of Rust code without reading it. He also polled the room on who enjoys receiving "slop grenades," and only a few hands went up.
He then turned serious to recognize Mike Dalessio. Patterson said they met more than twenty years ago, when Dalessio sent him a pull request for Mechanize. They have worked together in open source ever since, and at Shopify, where Dalessio was for a while Patterson's boss. The two run a GitHub organization called Sparkle Motion. In Patterson's joke, he is the "sparkle" who appears on stage with jazz hands, and Dalessio is the "motion" who does the work. Their shared projects include Nokogiri and the SQLite3 gem, and Patterson noted that Dalessio is very active on the Rails security team. He then announced that Dalessio is the newest member of the Rails core team.
What Patterson Worked On This Year
Patterson said he always talks about performance, and this year's topics were Ractors, ZJIT, and AI. He also mentioned a project he would not cover: he wrote a register allocator for ZJIT this year. He called it the hardest project he has ever worked on and said he was proud of it.
He added a bit about a new manager and their "first one-on-one." The jokes played on arrays/errors ("my code is exceptional") and on the "ensure" keyword ("these meetings will only get better. I'll ensure it."). The same routine led into the framing quote of the talk.
"The Purpose of a System Is What It Does" and Who Is Watching
Patterson said he had heard "the purpose of a system is what it does" many times and never liked it, because he found it too ambiguous. His example was a Nokia 3310. If he took one back in time and handed it to someone who had never seen a cell phone, that person might use it as a hammer. Read literally, the phrase would say the phone's purpose is to be a hammer. To Patterson, it is a phone.
He concluded that what a system does depends on who is observing it, and offered a revision: "the purpose of a system is what it does for those that are observing." He said this observational view is common among language implementers, where it is called the as-if rule. He read the C++ formulation: compilers may apply any optimizing transformation as long as the optimization makes no change in the program's observable behavior "as specified by the standard." Put simply, the compiler can change your code however it wants as long as it runs the way it looks like it should. He pointed out that Ruby has no standard comparable to C++'s, and set that problem aside.
For the rest of the talk, he examined optimizations as having two sides: what gets deleted, and who observes that deletion.
Memoization and the New Observers Ractors Bring
The first example was a pattern he expected everyone had written: memoizing an expensive computation into an instance variable. Patterson said this bakes in an assumption that only one thread observes the code at a time. There is already a race condition in which the computation could run more than once, but the GVL keeps developers from seeing it.
Ractors change that. They allow true parallelism, so two threads can hit the same code path at the same moment, and the set of observers changes. Patterson noted that Ractors raise an exception in this situation rather than silently corrupting state, which gives developers a chance to fix the code. He referred the audience to Andrew's talk earlier at the conference for more detail.
He did not want this to scare people away from memoization. He showed two cases. In the first, several Ractors access a memoized class-level instance variable on a class. That is a race, and the Ractor raises an exception. In the second, the pattern is identical but only one Ractor observes the code path, so it is fine. What matters is keeping the state isolated to a single Ractor.
Tools for Shared State: on_main, Lock Variables, and TVars
Class-level instance variables are common in codebases, and Patterson said the upcoming Rails release will offer an ActiveSupport Ractors module with an on_main method that runs specific code on the main Ractor. That lets an expensive computation run there and store its result in a class-level instance variable. As Andrew's talk explained, this adds latency. If much of the work is sent to the main Ractor, it becomes a bottleneck and every other Ractor contends for it.
Where possible, Patterson recommended a gem of Ractor-safe data structures written by Koichi Sasada, the main Ractor author. He pointed out two structures that address the memoization problem, and said the choice depends on the throughput you need:
- A lock variable guarantees a computation happens exactly once. Use it when the computation is not idempotent.
- A TVar allows the computation to run one or more times. Use it when the computation is idempotent and running it a few times is acceptable.
He said the gem has many more structures. He argued that because of how Ractors are implemented, these structures will not cause deadlocks or thread-safety bugs; misuse produces exceptions instead. He also argued the blast radius is small. His illustration was a hypothetical Ractor-based web server that reads a request from a socket in a loop and serves each request inside a Ractor. Most allocations in a request–response cycle happen inside the serve method. Code touches outside state such as constants and configuration values, but Patterson said he thinks those touch points are fewer than people expect.
Ractor-Local GC in Ruby 4.1 and Shared Objects
Next he looked at data observability and performance. Patterson said Ruby 4.1 will ship with a Ractor-local garbage collector. Today the GC is a singleton, so every Ractor that allocates has to ask the same collector, and Ractors running in parallel end up waiting on it. Ruby 4.1 gives each Ractor its own heap, which removes that bottleneck.
He suggested this might also allow cheaper collection. In the fake web server, you could create a new Ractor per request and, when the request finishes, abandon its heap. He said that theoretically this could make GC cheaper.
He then said the story is more complicated. In a producer–consumer example, the producer allocates an object in its own heap and passes it to a consumer. Both Ractors then point at the same object, which the matching object IDs confirmed. So when one Ractor dies, its heap cannot simply be thrown away, because other Ractors may still reference objects in it. Ruby reconciles this with a global GC. Patterson described it as similar to a major GC that hopefully runs less often than a major.
His second example showed the sharing can be implicit. He reused a demonstration from last year's Rails World: parsing two JSON documents yields keys that are the same object with the same object ID. When the parsing is split across Ractors, the "hello" key still has the same object ID in both, because keys are deduplicated across Ractors. Under the hood, that string is allocated in the main Ractor's heap and the other Ractors point into it. Patterson found this interesting because GC performance might be affected without anyone knowing, since it happens implicitly.
Callbacks and a Stack That Grows
Patterson then moved from optimizations developers write to optimizations in frameworks and the language, beginning with backtrace inconsistencies. He built a small fake ActiveRecord-style module using ActiveSupport callbacks and printed the call stack from inside a save callback. He expected three frames: the run_callbacks block, save, and main.
With no callbacks, he got three frames. With a before callback, still three. With an around callback, the stack trace "exploded," even though nothing inside save had changed.
He quoted a comment he found in ActiveSupport. It says that because the method wraps large portions of user code, it has "an additional design goal of minimizing its impact on the visible call stack," and that when control passes smoothly into the supplied block, "we want as little evidence as possible that we were here." Patterson said this is exactly the as-if rule. He forgave the extra frames with around callbacks, since the application really is doing more work.
Ruby 4.0's Missing Class#new Frame
He then asked whether a stack can show fewer frames than expected, joking that he spends weekends counting stack frames. His example chained methods one → two → three. three allocated a Foo, and Foo#initialize called four, which made the trace easy to read. The expected trace was initialize, new, three, two, one, main, and that is what Ruby 3.4 prints. On Ruby 4.0, the Class#new frame is gone.
To explain why, he dumped the VM instructions for three on both versions. Ruby 4.0 produces many more instructions, because the initialize call has been inlined at the new call site. Translated back into pseudo-Ruby, the logic is: if Foo.new is the default implementation, allocate a new Foo, call initialize directly, and return the object; otherwise take a slow path and call new normally. three now calls initialize itself, which is why the new frame disappears.
He showed the slow path by monkey patching Foo.new with a method that only calls super. Printing traces before and after the patch showed the fast path with no new frame and the slow path with Class#new back in the stack.
The benefit is speed. Patterson reported that this example allocates 70% faster on Ruby 4.0 than on the older version, without any JIT. The deleted thing was a call frame. Its observers are things like caller, debuggers, and other stack-inspecting tools. He called the trade worthwhile.
He said CRuby does this all the time. In another example, a class implementing hash prints a backtrace whenever it is hashed. Of three calls, he expected the first two to have identical traces and the third to differ but have the same depth. The traces were actually quite different, because the [] method's frame was elided in one case. CRuby "plays fast and loose with the stack trace," he said.
He asked the audience who cares, and only a few hands went up. That is the reaction he wants. If people did care whether those frames appeared, these features could not be built. He called such cases gray areas of the as-if rule: execution is not strictly the same and the difference is observable, but nobody minds. A language has to preserve behavior where it matters, and Patterson called optimizations that trade away behavior nobody cares about "low-risk optimizations."
ZJIT's High-Risk Bets and Patch Points
Next came "high-risk optimizations" in ZJIT. Patterson said ZJIT can make them because it can deoptimize when something unexpected happens. ZJIT has two intermediate representations: a high-level IR (HIR) and a low-level IR (LIR). The HIR can be inspected by passing a ZJIT dump-HIR flag to the Ruby binary.
One bet ZJIT makes is that you probably won't redefine methods, but if you do, it has to notice and behave correctly. His example allocated a point, then called a method that called a method on the point, inside a 5.times block. The HIR showed that get had been inlined into the block with an inline frame pushed for it, and that the constant 42 had been lifted into the block. ZJIT had reduced the code on the left of his slide to much simpler code on the right.
Patterson said that while people might not care about a missing stack frame, they would care if their code redefined get or x and the JIT ignored it, since that would break the as-if rule. The HIR contains patch points, one marking possible redefinition of get and another for x. Patch points emit no machine code. They are markers for where the machine code can later be patched.
He showed machine code before and after invalidation. When the method is redefined, the JIT overwrites a test instruction at the invalidation marker with a jump to exit code. The exit code restores the VM's internal state and stack, and execution continues in Ruby's interpreter.
Detection relies on CRuby. When a method is redefined, CRuby already has to invalidate caches, and that same point is where the JIT can invalidate compiled code. The JIT tracks which methods it compiled and throws away the associated machine code when one is redefined, for YJIT or ZJIT. Method redefinition is only one invariant. ZJIT also watches basic operations ("bops"), making sure + still means +. It watches constants for redefinition, and it watches TracePoints, because a TracePoint could observe things the JIT optimized away. He said there are others and invited questions after the talk.
"AI" in ZJIT: Abstract Interpretation and Allocation Elimination
Patterson joked that since this was "AI world," he should mention that ZJIT uses a lot of AI, meaning abstract interpretation. He invited anyone interested in "AI" to contribute to ZJIT.
He defined abstract interpretation as pretending to run code and letting optimizations fall out of that. His demo was a test server and a small Rack application. A helper counted objects allocated while serving about 1,000 fake requests. Without a JIT, the run allocated about 1,000 objects, one per request. YJIT allocated about the same. ZJIT allocated eight.
The source of the allocations was clear: the Rack call method returns an array of status, headers, and body. ZJIT inlines call into the server's serve method. Patterson then rewrote the inlined code into small steps with temporary variables, which behaves the same, and walked through what the abstract interpreter does with it.
The interpreter builds an abstract heap that tracks what objects would look like if the program ran. Stepping through, v1 is 200, v2 is the headers constant, and v3 is the body. v4 is an abstract array containing v1, v2, and v3; the interpreter doesn't care what those values are, only that the array holds them. status, headers, and body are then assigned from the array's elements. By algebraic substitution, status is simply v1, and likewise for the others. After substitution, v4 is never used, so it can be removed. The intermediate variables can be removed too, leaving direct assignments. Repeating this with more inlining and constant folding is how ZJIT removed the allocations.
He then asked the audience whether this violates the as-if rule, since counting allocations can detect the change. He said he personally doesn't care and thinks reducing allocations is great.
The RubyGems Incident: A Timeline
Patterson said he had intended to end there, and asked for applause for the eliminated objects, but he had a second story. He told it from his own perspective and said he had built the slides the day before.
On May 11 and 12, RubyGems.org began receiving thousands of junk gems. The campaign was later called "GemStuffer." On May 12, Maciej Mensfeld, who works with RubyGems.org and scans new packages, disclosed the attack on Twitter, saying they were dealing with a major malicious attack and that signups were paused. On May 13, Socket published a blog post naming it the GemStuffer campaign. According to that post, the gems contained a script that downloaded data from UK government sites, packaged it into a new gem, and uploaded that gem to RubyGems.org.
Patterson said the post confused him. He knew that installing a gem with a C extension runs its extconf, which is a known remote code execution vector. But the malicious script was not in an extconf, so installing the gem would not run it. He said the article never explained how or when the script would execute. He moved on, and registration reopened on May 16.
The Caching Vulnerability
On July 6, Luke Marshall of Truffle Security reported a caching problem to the RubyGems.org team. Patterson noted that he works on the RubyGems client team, not the server team, though the two collaborate.
Legacy authorization used a GET request. A client such as gem authenticating with basic auth received a response containing the user's API key, and Patterson asked the audience to remember the key's format. Fastly sat in front of RubyGems.org and cached GET responses. An attacker could send the same request without an authorization header and receive a cached valid key. Patterson said the bug had been in production for roughly six years.
He offered several reasons it went unnoticed. There was no known evidence of abuse at the time. It only triggered with Accept-Encoding: gzip, which curl does not send by default, though Net::HTTP does. And from a maintainer's point of view, receiving someone else's key would just cause a confusing "not authorized" error on gem push, since you would be pushing your own gem with the wrong key. Running gem login would fix it and you would forget about it. Only old RubyGems client code used this path, and the cache expired after an hour. RubyGems shipped a complete fix three days after the report.
An Email Chain and a Gem Called sln_leaker_5
On September 8, Patterson's colleague Emily told him to expect an email from Gemma, a former coworker now at Anthropic, about a serious RubyGems hack. Emily messaged him first because, as Patterson admitted, he is bad at email. Gemma introduced him to someone named Neve. Patterson panicked, searching his inbox and HackerOne for a report he might have missed, and found nothing. Neve replied that it was not a current RubyGems problem; a security researcher was trying to reach someone on the RubyGems team.
That researcher, Sydney von Arcs, told him that OpenAI agents had attempted to exploit the caching vulnerability, and that four package submissions appeared to try to steal users' API keys. Patterson was skeptical until he read the linked code. He pointed out a comment at the top referring to leaked key variants, the file name script.rb, and the gem name sln_leaker_5. The script sent a GET request to the RubyGems.org authorization path, and at the bottom was a regular expression matching the API key format from the auth response. He concluded it was indeed trying to hack RubyGems.org.
How rubydoc.info Became the Execution Engine
In a meeting with the researchers, Patterson asked how script.rb was ever executed, since it wasn't in an extconf. Their answer was rubydoc.info, a documentation server for gems that is separate from RubyGems.org.
The gems included a .yardopts file that loaded script.rb. When the YARD gem is installed, it installs a RubyGems plugin. When another gem is installed afterward, YARD looks for a .yardopts file in it and runs the code referenced there.
The resulting loop worked like this:
- A bot uploaded a gem to RubyGems.org.
- RubyGems.org sent a webhook to rubydoc.info.
- rubydoc.info downloaded the gem and opened it inside a Docker container that still had network access.
- The
.yardoptsfile caused the malicious script to run. - The script downloaded data from UK government sites, packaged it as a gem, and uploaded it to RubyGems.org.
- The cycle started again.
Patterson called it "insane" and a "Rube Goldberg machine." He stressed the timeline. The GemStuffer gems were uploaded around May 11. Truffle Security's independent cache report, which had no affiliation with OpenAI, came on July 6. By his reading, the bots knew about the cache bug a month or two before it was reported. He added that this all happened well before the Hugging Face breach. On September 12, Sydney's team published a detailed write-up at rubyhack.ai.
He listed what unsettled him: the bots worked out that RubyGems.org sends webhooks to rubydoc.info, that YARD executes arbitrary code, and that rubydoc.info processes YARD options with network access. He also offered a theory, which he labeled as his suspicion. Not every GemStuffer gem contains the key-stealing code. He thinks that when the RubyGems.org team started shutting down the bots' accounts, the bots needed a way to authenticate without their own credentials, and "decided to burn a zero day on rubygems.org."
He said the story ended well. The RubyGems.org team handled it completely, closed the leaks, and as far as anyone knows, nobody was exploited. He thanked them.
An Optimist, Not a Surrenderer: AI Is Not Your Compiler
Patterson said he uses AI every day to write code, then corrected himself: he is an optimist in general, not specifically an AI optimist. He believes this is the best time to be alive and that things will get better. But he said being optimistic about AI and the future does not mean surrendering to it.
He addressed the argument that AI is just the next compiler: you don't read your compiler's machine code, so why read the AI's code? He said he partly accepts it, since people really don't read compiler output. But he argued that is because the compiler made a promise, the as-if rule: whatever it does, the behavior you can observe matches the code you wrote. Everything in his talk, from Ractor exceptions and callback stack hygiene to deoptimization via patch points and careful allocation elimination, showed what it costs to keep that promise.
"Your AI has not made this promise to you," he said. "There is no as-if rule for your English." He said he would not concede his destiny to AI and would keep reading his code, and hoped the audience would too. His last slide read: "Get in, loser. We're going programming."
Hello everybody. Hello. Hello. Yes. Before we get started here, I want to make a quick announcement. I've started my own Linux distribution.
I wanted to share with you booting it up, but I couldn't put a machine on stage. So, I sent it back with the AV folks, so I'm going to have them start up my presentation now, please. Okay. All right. Okay. They're going to boot it up.
Okay. When it... Can you open PowerPoint for me when it... I vibe coded this, by the way.
All right. Okay. So I guess let's get going. I'm calling it Aaron XP, the name of this software. It installs in 100 milliseconds. And it's very important for me to get the install time down because I'm doing that constantly. Wake up in the morning, install my OS, get some coffee, install my OS, upgrade Chrome, reinstall the OS. I mean, come on.
I think it's funny, I think it's ironic to make a joke about the financial crisis and then encourage everybody to generate a whole bunch of Rust code without reading it. I'm going to start selling insurance on technical debt. So come see me afterwards if you'd like to buy some.
I know everybody here loves using AI. I do too. I love using AI. Love using it to make slop grenades. It's great. Really good. But I want to take a poll of the audience here. Yeah, I'm going to take a poll. Okay. How many of you like receiving slop grenades? Show hands here. Okay, we got two, three. Ah, a couple dedicated folks in the front here. Great. Thank you. Nice to see this.
I also got some... Thank you, Amanda, so much for your hard work here. Can we give her a round of applause? Yes.
So, when she was showing off the records during this conference, somebody got a photo of me here in the audience. They said, "This is me."
I had to tell her this afterwards and then I was like, "Now you can't... I'm making this joke. Now you can't say it on stage. It's mine. I'm doing this."
Okay. Before I make any more jokes today, I want to do something a little bit serious. I'm going to call out somebody here in the audience, a specific person. I'm going to embarrass them, and he is right there. Mike Dalessio. So I don't know if you saw his talk, but Mike posted this. We met over 20 years ago now. He sent an open source pull request to me for Mechanize, a Ruby library at the time. And I have been working with him in the open source community ever since. I even was able to work with him at Shopify for a while. He was actually my boss there.
We started an organization on GitHub called Sparkle Motion. I am Sparkle, he is Motion in this relationship. What this means is I come up on stage with jazz hands and he actually does work. We work together on many, many open source projects. For example Nokogiri, maybe you know it. I have seen it in various slides about slow gem installations.
We've also worked on the SQLite3 gem, and he's also very, very active in the Rails security team. But the announcement that I want to make here today is that he is also the newest member of the Rails core team.
So, thank you. Thank you, Mike. Really, really, really appreciate your work. I don't know if the website's updated or not yet, but anyway, thank you. I'm looking forward... Thank you so much, and I'm glad you will continue to be active supporting us. Yes.
All right, it's exciting to be here at Rails World, and I know many people have made this joke before, but I am also using Claude to generate my slides. So, it's great to be here at AI World in Texas. It was cool to learn about how we can use AI. All of us, we can all use AI to become thousandx programmers. So, all of us here today, we can use AI to become thousandx programmers. Sorry, hold on. I got a typo. I mean, thousandx programmers.
Anyone here? Did anyone here ride the robotaxis? This is amazing to me. Yes. Yes. If you haven't ridden them yet, I highly suggest that you do. It's really, really fun. But make sure to bring a friend with you because it's way more fun that way. Yes. Okay.
Happy Friday. Happy Friday, everyone. It's always Friday somewhere. My name is Aaron Patterson. I'm also known as Tenderlove. I work on the Ruby and Rails infrastructure team at Shopify, and I help with various things around the team. I work on a lot of stuff there. It is very fun. I also have a cat. This is my cat. You may have seen him here at the check-in at the passport booth.
Today I'm going to be talking about performance, obviously, because that's what I always talk about at presentations. I don't know what else to talk about. I'm going to talk about Ractors. I'm going to talk about ZJIT. And of course, I am going to be talking about AI because we are at AI World. So we will be doing that.
And the reason I'm going to be talking about these things, the important thing, is because this year I have been working on performance and Ractors and ZJIT and also AI. And yes, I am actually repeating these slides because I get very, very nervous when I'm preparing presentations that I will not have enough content for the conference. So I will just add more slides. But yes, this is what I'm going to talk about today.
This year I also wrote a register allocator for ZJIT. And we're not going to talk about it today, but I wanted to put this in here because this is the hardest project I've ever worked on. And I'm very, very proud of my work here. And I just wanted to put this here and say I did that and it was great. So I am proud of myself for doing this. Oh, thank you.
I also got a new manager this year, and things are going pretty well with him, I think. I'm not quite sure. I recorded my first one-on-one with him and I'm going to share the transcript with you today, and hopefully you can help me be the judge of this. So, in our very first one-on-one, he said to me, "Aaron, you've only been testing code paths that make errors." And I said, "Oh, well, that is because I am expecting arrays. It's because my code is exceptional."
And he was not very pleased with this, I don't think, but I wanted to reassure him. So I said to him, "Don't worry. These meetings will only get better. I'll ensure it." Yes. So I think things are going well. I'm not sure.
I'm going to start my presentation with a quote here: "The purpose of a system is what it does." I've heard this quote many times before, and I never really thought about what it meant too much. I don't know. It makes me think and I don't like doing that, so I didn't really think about it too much. But I guess I brought it up with my manager, and I thought to myself, okay, when I'm trying to understand what a system is for, I need to understand what it does.
And since my manager was new here, he said to me, "Aaron, what is this system for here that you're working on?" And I told him, "Well, it is... what does the system do?" And he told me to go half myself. I'm kidding. I'm very much kidding here. He would not say that to me. I said its purpose is to raise exceptions. So, I'm not sure. I think these one-on-ones are going well. I don't know. We'll see after this.
Anyway, I don't particularly like this phrase. And the reason I don't like it is because it's so ambiguous. I think about it, and immediately I'm like, what does this system do? What does it do? I don't know. So I don't know what the purpose is. I can't figure it out. My brain is just way too small for this. I can't deal.
To give you an example, let's say I traveled back in time with my trusty Nokia 3310. Yes. Anyone have this phone? Yeah. Heck yeah. I'm sure there's people in here that have never seen this before, but this is an indestructible phone. It is great. Let's say I traveled back in time with this phone and I gave it to someone that had never seen cell phones before. They may actually use this as a hammer.
So then I wonder, is the purpose of this phone to be a hammer? If you took that phrase that I put before literally, you might say yes. I mean, clearly they're using it as a hammer. So that is the purpose of this thing. But to me it would be a cell phone.
I think what the system does depends on who is actually looking at the system. So when I'm looking at that phone, I see a phone. I will use it as a phone. But when I brought it back in time, the person that sees it there, they see it as a hammer and they will use it as a hammer. So really what it depends on is who is making the observation of that system. So, if I could refine the original quote, I would say the purpose of a system is what it does for those that are observing. Somebody famous said that at Rails World 2026. Yes.
So, I've been thinking about this a lot because, as I said, I've mostly been working on performance optimizations or code optimizations. And this particular attitude that I'm telling you, this observational attitude, is commonly known among language implementers. It's known as the as-if rule. There's a rule called the as-if rule, and you can go read about it on Wikipedia, but this is actually codified in the C++ programming language with regard to applying optimizations to code.
So, we're going to read an excerpt from the as-if rule. The as-if rule says in C++ they are allowed to apply any optimizing transformation to a program during compiling, provided that such optimization makes no change in the observable behavior of the program as specified by the standard. So in other words, the compiler is free to make any changes to your code that it wants to, as long as the code executes the way that you wrote it, the way that it looks like it should execute.
Now there is kind of an awkward thing in this particular quote. It says "as specified by the standard," which we pretty much don't have in Ruby. There isn't a specific standard like C++ has. But let's not worry about that for now.
So today I want to talk about optimizations through the lens of observable behavior. Every optimization has two sides to it: the thing that gets deleted in the optimization, and the thing that observes that particular optimization, whoever observes that code. We're going to look at some examples of these different observers and optimizations and see how they impact each other.
So the first one that I'm going to show you, I guarantee pretty much all of you have done this before, I think, is memoization. This is the very first example. I'm sure we've all written this particular pattern before. We have some function. We know that that function is expensive, but we only need to calculate it once. So, we calculate it once and then save it off into an instance variable. I'm pretty sure many folks here have written code similar to this. I know that I have.
The only problem with this particular code pattern is that we've baked an assumption into the code. And the assumption is that this code is only ever observed by one thread at one time. So there's actually a race condition in this code where it could end up that we execute this expensive computation multiple times. But we don't really see that because we have a GVL in Ruby.
Now this is kind of where Ractors come in. Ractors allow us to have true parallel processing in our Ruby code. That means that the observers of our code have changed. We can actually have two threads see this exact code path at the same time. So that means we can have multiple observers of a particular code path in parallel.
Now fortunately, Ractors will actually raise an exception for us in this case. So instead of having a race condition and messing up something in your application, you'll actually get an exception and you have the opportunity to fix it. Now hopefully you saw Andrew's presentation. He talked in detail about this type of stuff, but I'm going to repeat a little bit of it here.
So, when I say this though, I don't want this to scare you off and say, "Well, I should never use this particular pattern." I'm going to show you a couple examples of the pattern. In the first case, we have multiple Ractors that are accessing this memoized instance variable on the Foo class. And what that means is that we have a race condition on the Foo class. And the nice thing is that this R1 Ractor will actually give us an exception saying, hey, you can't do that, you need to do this in a thread-safe manner.
Now the second case is actually fine. You can see that the pattern is the same. We're able to use it. It's really, really fine to do this. It's just that in the second case we only have one observer of that particular code path. So it's okay to use this pattern. We just need to make sure that it's isolated to one specific Ractor.
Now, fortunately, this top pattern of class-level instance variables, we see this fairly commonly throughout code bases, and there's actually going to be a way for us to mitigate or deal with this situation in the upcoming Rails release, and that's with the Active Support Ractors module. With the Active Support Ractors module, there's an on_main method that allows us to execute certain code on the main Ractor. So if we have an expensive computation and we need that to occur inside of the main Ractor and stash something into a class-level instance variable, we can do that.
Now if you saw Andrew's presentation, you'll know that this causes latency in your application. And the reason it causes latency is because if you shuffle lots of computation off to the main Ractor, all of a sudden the main Ractor becomes the bottleneck for your application, and you'll end up seeing increased latency because everybody is contending for that Ractor.
So when possible, I highly recommend using the Ractor sharing gem. It's written by Koichi Sasada, and he's the main Ractor author. So hopefully this will work pretty well with Ractor-based code. Actually, not hopefully, it will. So you should use this. It has many different Ractor-safe data structures inside of it. I'm going to share a couple of them. I'm going to Ractor-share a couple of them here.
It's got two here that I want to point out that solve this particular problem that we were looking at. First is lock variables and the second one is TVars, and the one that you choose depends on the throughput that you want with your application. A lock variable will guarantee that a particular computation only occurs once, and the TVar allows it to occur one or more, a few times basically. So the first one you use when you have a computation which is not idempotent and you need to guarantee it only occurs once, and the second one is when you have a calculation that is idempotent and maybe it's fine if it does it a couple times. Plus there are many more data structures in there.
And I don't want to scare you away from using these data structures, because I think that the nature of the way that Ractors are implemented means that even if we use these data structures, there won't be any deadlocking in our applications and there won't be any thread safety issues in our applications, because we'll just get exceptions instead.
And the reason I think we shouldn't be too scared of these data structures is that I think that the blast radius is fairly limited. So let's take an example here. We have a pretend Ractor-based web server. So it just reads a request from the socket in an infinite loop, and then it serves up the request inside of a Ractor. And if you think about the normal request-response lifetime of this, we're allocating most of our objects inside of this serve method. So all of our allocations will happen inside of there. And it's really
The only time we really need to worry about this is when we touch things on the outside. Now of course we do touch things on the outside, like constants or configuration values, things of that nature, but I think that the amount of touch points is much lower than we actually think.
The other thing I want to talk a little bit about is data observability and how that impacts performance. Ruby 4.1 is going to be shipping with a Ractor-local garbage collector. So what does this mean? The GC in Ruby is a singleton. So whenever a Ractor needs to allocate an object, it's going to go ask the garbage collector, hey, give me an object. And if we have multiple Ractors running in parallel, well, it's pretty clear that we're going to run into a bottleneck here. So we'll try to allocate a bunch of stuff and then these Ractors are all waiting on the GC to allocate things. So in Ruby 4.1 we'll have a Ractor-local garbage collector, which means separate heaps for each Ractor, which will eliminate this particular bottleneck. So each Ractor can go ahead and ask the heap for an object.
Now we might also be able to leverage this to have cheaper garbage collection. So for example, in our fake web server here, we can create a new Ractor on every single request and then when the request is done we can just abandon the heap and then continue the process. So theoretically we can get cheaper GCs out of this.
Now I wish that was the end of the story, but the story is a bit more complicated than this and I want to show you an example of why. Let's say we have this producer consumer example in Ruby. I apologize for making you read code. The producer allocates an object in its own heap like this and then it passes that object to a consumer. So when that happens, the consumer and the producer are both pointing at exactly the same object. And you can see that with the object ID here. The two Ractors are able to observe exactly the same object at the same time.
So what that means is that when one of these Ractors dies, that doesn't necessarily mean we can throw away the heap, because it could mean that there are others that are pointing into that heap. So we need some way to reconcile this situation. The way we reconcile it is with a global garbage collector, a global GC. And you can think of this kind of as a major, except that it hopefully occurs less frequently than a major.
So I want to show one more example. In the previous example, we purposely passed an object from one Ractor to another and they were both able to see that object. And I want to show a different example where we're doing it implicitly. So I showed this example at Rails World last year, but I think it's a great example. We have two JSON parses. We're parsing two JSON documents and then looking at the key in the JSON document. And you'll see here that those keys are exactly the same object. They have the same object ID. We're able to observe they are the same object.
Now, if we take this code and split it up into multiple Ractors, we'll see that that string hello, it's the same object ID in both Ractors. So in both Ractors, those keys are deduplicated among the Ractors, and both Ractors see the same hello string. And what's happening under the hood is that we're allocating a hello string in one heap, our main Ractor heap, and then the other two Ractors are pointing into that heap. So what's interesting about this example is that it possibly means we're impacting GC performance, but we don't know it. It's happening implicitly.
So I want to take a step back for a minute. We've been looking at stuff a little bit deep here. We started this by looking at optimizations that developers do. So we looked at how Ractors add a new observer to the system and the downstream impacts of adding new observers. Now I want to look at the framework and language optimizations, and specifically I want to look at backtrace inconsistencies.
So here we have an example of a save callback. We're using Active Support callbacks. We're making kind of a fake Active Record thing here where you can mix this in and then call save and then you end up with a save callback. And this example code prints the call stack from inside of our save callback. If we imagine the stack trace from this, we can imagine that first we'll see the block for run_callbacks, then we're going to see the frame for the save method, and since we're running this just from a script, we'll see main as the top level.
So, let's take this module and we'll mix it into a class, a no callbacks class. So, this is kind of our base case here. If we run that, we'll see indeed, yes, it only has three stack frames. So, let's try it again, but we'll do it with one callback. In this case, we'll set a before callback and look at the backtrace. And we see that again there's only three frames here. So that's great and it's what we expected. Now we're going to do one more. We're going to do an around callback. If we do that, we'll see that our stack trace has all of a sudden exploded. Now remember, we didn't change any of the code that was inside of that save method, and all of a sudden, our stack trace is wildly different.
Now, as I was looking into this, I thought it was very funny to find this comment in Active Support. I'm going to read it. It says: "As this method is used in many places and often wraps large portions of user code, it has an additional design goal of minimizing its impact on the visible call stack. An exception from inside a before or after callback can be as noisy as it likes. But when control is passed smoothly into the supplied block, we want as little evidence as possible that we were here." And this is exactly the as-if rule that we were talking about earlier.
So in this case we were seeing more stack than expected, and this was happening inside of Rails, and I can forgive that because we have application code that is doing more stuff. We added in an around callback, so we kind of know, yeah, okay, it's doing more work, so of course there's going to be more stack trace. But is it possible to get less stack trace than expected? Do we see that case? We do. But for us, yes, I do like to spend my weekends counting stack frames. I don't think that makes me a loser, though.
So here we have an example of a missing new stack frame. I'm going to show what this is. So in this code we have a method one that calls method two, then calls three, which allocates Foo. Foo calls four from inside of initialize. And the reason I'm doing this is just to make the stack trace a little bit more clear. So if we run this code, we expect to see a stack trace like this. We'll see initialize, new, three, two, one, main. Nothing surprising going on.
If we run the code in Ruby 3.4, that's what we see. This is exactly the stack trace that we see in Ruby 3.4. If we run this in Ruby 4.0, we'll see it looks a little bit different. And in fact, in Ruby 4.0, we are missing this Class#new frame. It's gone. So, what happened to this stack frame? I think we can find some clues. We're going to figure it out here.
If we dump the instructions for the three method, I think we'll be able to learn something. So, we're going to dump the VM instructions for this particular method. On the top, we have the instruction sequences for Ruby 3.4 and on the bottom, we have the instruction sequences for Ruby 4.0. And the first thing that we can notice from this is that Ruby 4.0 generates a lot more instructions than Ruby 3.4. I'm not going to explain these instructions fully, but the basic change is that we've inlined the initialize call into the new call site, right there. We did an inline.
So if we take those instructions and kind of write them as Ruby code again, on the left here is the code that we wrote and on the right is kind of the pseudo code of the instructions that we generated. The pseudo code of the instructions is basically saying, hey, if the new method on Foo is the default implementation, then what we're going to do is we're going to allocate a new Foo and then we're going to directly call initialize and then we're going to return the object. If it's not the default new, then we're going to take a slow path and just call the new method.
So it makes sense that there are more instructions here because it's just doing more than Ruby 3.4 was doing. The other thing to notice here is that we're actually calling the initialize method from the three method. So before, the model was that we would call new, new would call initialize, and now we can see here three is directly calling initialize. So it makes sense why that new stack frame is missing.
So let's say we have this code. It's our Foo class. We print out the stack trace from initialize before implementing new. So we're going to monkey patch new basically. But it's not like a big deal monkey patch. We're just calling super from inside of it. So here we'll print out the stack trace before, we'll add new, and then we'll call the method again and print out the stack trace again. If we do that, on the left is our stack trace from the fast path and on the right is our stack trace from the slow path. And we can see that our call to Class#new is back.
So why did we go through all this trouble? Why do all of this stuff? The reason is because we get faster allocations out of it. This particular example here allocates 70% faster on Ruby 4.0 than it does on Ruby 3.0, and that's without a JIT compiler. So we get really good performance gains out of this particular optimization.
So the thing that we deleted in this case was a call frame, and the observers of this would be like caller, debuggers, things that might inspect the stack. So these are the things that we impacted. The thing that we got out of it, though, is we got 70% faster allocations, which I think is a good exchange.
And CRuby does this all the time. There are many different places where it does it. I'm going to give another example here really quickly. Here we have an example Foo class that implements hash. It prints the stack trace whenever something calls hash on it. We trigger a hash computation, it prints a stack trace. So we'd expect these two first calls to have the same stack trace and we'd expect the last one to have a different stack trace but the same depth. And if we look at it, the stacks are completely different. Here we're able to elide the square square method in the very left hand side. So you'll see that that frame is missing, whereas the other two sides show you everything. So CRuby does this all the time, and it plays fast and loose with the stack trace.
But I think the question is, who cares? Really? Does anybody care about this? Anyone? Ah, a few people. Okay, I know. I think it's interesting. But actually, that is what I want you to feel. If you actually cared whether or not those things were in the stack frame, then we wouldn't be able to implement these kinds of features. So that's why I'm very, very glad that most of you are like, I don't care about this.
I think these types of things are as-if rule gray areas. So the compiler made a change to the code, so execution strictly isn't the same. You can observe a difference, but these cases are cases where people don't care. They really don't care, so it's fine. And in Ruby, and I think most other languages, we need to maintain behavior where it actually matters. So those cases don't actually matter. We're getting performance optimizations and trading for things that don't really matter. So I think these types of optimizations are what I like to call low-risk optimizations, where we change the language behavior but people in general really don't care. They're low risk.
So next I want to look at some high-risk optimizations, specifically inside of ZJIT. ZJIT is able to make high-risk optimizations because it's able to deoptimize in the case that something unexpected happens. So we're going to look at some of those. ZJIT has two types of intermediate representation internally. We're going to look at these. An intermediate representation, also called IR, has a high-level IR, or HIR, and a low-level IR, or LIR. If you want to try this out at home, you can see the HIR by adding these flags to your Ruby binary. So if you add zjit-dump-hir, you can see the output and take a look at it.
One of the bets that we make in ZJIT is that you probably won't redefine methods. Probably not going to do it. But if you do do it, we need to know about it, because if you do it, we need to behave correctly. So let's say we have an example like this, where we allocate a point, then call a method which calls a method on the point.
If we dump the HIR for this, it'll look like this, and I'm going to highlight a few things inside of this HIR. First I want to show off we have a block here. So we can see that this is the HIR for the block in 5.times. The next thing I want to point out is that we've inlined get into the block, but we push an inline frame. So, we're pushing an inline frame for get. So, it's inlined into the block. The next thing I want to show is that we've actually been able to lift this constant 42 down inside of the block. So, ZJIT was able to take the code on the left and compile it down to the code on the right. It was able to simplify all of that.
But how can ZJIT make such big bets? We might not care if a stack frame is missing, but if your code redefines get or x, you would probably care, because now your code isn't behaving the way that you expected it to. It's not following the as-if rule. So if we go back to the HIR output, we'll see a couple of PatchPoint lines. I'm going to highlight those. We have one here. There's a patch point for method redefined on get. And we have another patch point defined here for a method redefinition on x.
And these patch points don't represent any code. So if you look at the machine code output, there is nothing in the machine code that represents these patch points. They're simply markers. We use these markers to know where in the machine code to patch.
So let's say we have some code like our example from before. In this case, we're going to JIT compile the 5.times block. Then we're going to redefine the x method and then JIT compile it again. So if we do that, here's an excerpt from the machine code that the JIT generates, and we're going to show before and after invalidation. So this is before invalidation. So when the get method is invalidated, the JIT compiler will literally overwrite the machine code. So it'll take that test code. We have an invalidation marker that points at the test instruction. And when that method gets redefined, we'll actually just overwrite that test instruction with a jump instruction. And that jump instruction will jump to exit code. That exit code will restore all of the VM internals and VM stack, and you'll continue executing your code back on Ruby's virtual machine.
So how do we detect this case? How do we know that you redefined a method? The way we do it is with help from the CRuby implementation. Anytime a method gets redefined, we need to invalidate caches anyway. And this is also the perfect point for the JIT compiler to invalidate any JIT code. So the JIT compiler can simply track which methods it compiled. And if somebody redefines a method, then it can invalidate any machine code associated with that method. So here we'll say, oh, a method was redefined. We're going to invalidate all of our JIT code for YJIT or ZJIT, depending on which one you're using.
So method redefinition isn't the only invariant that we track. We also track things we call basic operations, or BOPs. So for example the plus method, we need to ensure that plus still means plus. We track constants, so we need to make sure that you don't redefine constants, or that there are
trace points, because if you run a trace point it means that you can observe things that the JIT compiler may have optimized out, and we need to fix that. It won't follow the as-if rule. And there are a bunch of other ones too. And if you want to know about the other ones, please come talk to me afterwards.
And as I said, I would talk about AI because this is AI world, and ZJIT uses a lot of AI. We use it a lot. And unfortunately, I know this is going to disappoint some people, but I mean abstract interpretation. Yay. So if you want to work on some AI, come work on ZJIT please. We would love to have the help.
So what is abstract interpretation? Abstract interpretation is when we pretend to run some code. We're going to pretend to run code and then allow optimizations to fall out from that. So I'm going to show an example to help motivate what this exactly means. On the left here, we have a test server. This is just a test Rack application. On the right is our actual, or excuse me, a test server. On the right is our Rack application. And then down here on the bottom we just have a simple method that measures the number of objects that we've allocated, and it tries to serve up about a thousand fake requests through this Rack application.
If we run this with a JIT compiler, or if we run this without a JIT compiler, it'll allocate about a thousand objects. That makes sense because we made about a thousand requests. If we run it with YJIT, we'll see it allocates about the same number of objects, and then if we run this with ZJIT, we'll see that it allocates eight objects. So where are these objects coming from, and how is ZJIT able to remove them?
So let's go back to our demo program, and it's pretty clear from the demo program where these allocations are coming from: allocating an array here in the call method with our status, headers and body. So we already saw from the previous example that ZJIT can inline functions. So what's going on here is ZJIT is taking that call method and inlining it inside of the serve method. So let's say we did that. We're going to inline it here like this.
Now what I want to do here is, let's say we take this code and we break it down into more simple pieces. We're going to break it into multiple steps with temporary variables and do this a piece at a time. So, we're going to take this status, headers and body assignment and break it up. So, if we did that, we would end up with something like this, where we've assigned all of these variables to intermediate variables and then we use them like this. And I think we can all agree here that the behavior of this program is the same as the previous one. It's just that we have more temporary variables.
And here is where abstract interpretation comes in. Our abstract interpreter will walk over this code and pretend to execute it. So we're going to walk through that now. So the first thing we'll do is we'll create an abstract heap. And this abstract heap keeps track of what objects would look like if we were to run the program, and we simulate this execution. So we create an abstract heap and then store the values in it. So we're going to simulate execution. We know if we execute the first line, V1 is going to be assigned to 200. V2 is going to be assigned to the constant headers. V3 will be assigned to the body. And then something special happens at V4. V4 is going to be assigned to an abstract array. So we have an abstract array here. This abstract array contains three values: V1, V2, and V3. And we don't care what V1, V2, and V3 are. We don't care what they are at all. We just know that it has those values inside of it.
After that, we're going to assign status to the first element of the array. Then we're going to assign headers to the second element of the array. Then body to the third element of the array. Now, the next thing that we can do is we say, well, we're going to perform a substitution here. We know that status is assigned to the value V1. So, we can just do an algebraic substitution in this case, and we'll say okay, we're just going to substitute those values out. And if we do that, there's something that we notice in this program, and that is that V4 is not used. We don't use V4 whatsoever. Since we don't use that, it means that we're allowed to eliminate it. So we eliminate that. Then we go even further and we say, well, you know, we don't need these intermediate variables either, so we'll just move those down to status, headers and body and do direct assignments. And we can be more aggressive. We can say, well, let's go down here as well, so we'll eliminate all those, and we can inline other functions and perform the same type of algorithm over and over again.
And in this way, using inlining and constant folding and abstract interpretation, we're able to convert the code on the left into the code on the right, and that's how we are able to eliminate those allocations.
So we talked about optimizations and the impacts that observers have on these optimizations. We talked about parallelism with Ractors, stack frame eliminations in the interpreter, and inlining and allocation elimination in the JIT compiler. So we are able to eliminate all those allocations in this JIT compiler. And I have a question for all of you in the audience. Does this optimization violate the as-if rule? We are able to observe that change by counting the number of allocations in the code. So do we violate the as-if rule? Do you care? I particularly don't. I think it's great that we're reducing those.
All right. I was gonna end the presentation here. So I was going to say thank you. We did it. We removed some allocations. But I actually have another story to tell. I apologize. There's actually two presentations in one. We're going to move on to the second one here. I've got 12 minutes left. We're going to try and do it. I want to talk about something that happened very, very recently to us in the community, and the story that came out from it. I think it's very interesting. I hope it is interesting to all of you. I want to talk about RubyGems and the OpenAI incident.
Did anyone hear about this? Few people. Yes. Okay. So I hope that this story is interesting, and for those of you that haven't heard the story, I hope it is still interesting.
On May... I want to recount this story from my perspective. Basically, on May 11th and 12th there was something that happened. rubygems.org started getting thousands of junk gems being uploaded. And actually, I want to rewind a little bit. Let me rewind here. I was going to end here. Thank you everybody. Let's have a round of applause for me. All those objects eliminated. Great. All right. Get the energy up a little bit.
So on May 11th and 12th, rubygems.org started getting thousands of junk email, or junk gems, uploaded. Just tons of them. Later on, this would be called the Gem Stuffer campaign, but there wasn't any name for it in particular. I mean, when you're getting a bunch uploaded, you don't really take time to come up with a name.
On May 12th, Moshe Mensfield disclosed the attack on Twitter. So he works with rubygems.org and scans new packages that get uploaded. They were getting tons of packages uploaded. They said, "All right, we got to stop the bleeding here. We're going to shut this down. We are dealing with a major malicious attack on rubygems.org right now. Signups are paused." So he said, "We're dealing with a major malicious attack on rubygems.org," and announced that signups had been paused for the time being.
On May 13th, soak.dev released a blog post calling this the Gem Stuffer attack, and at the time I was like, okay, that's interesting. I read their blog post about it. So their blog post is here and you can scan the QR code here. They called it the Gem Stuffer attack, Gem Stuffer campaign. So, I read the blog post, and the blog post was like, okay, there are these gems that are being uploaded, and the gems have a malicious script in them, and the script will download some stuff from the UK government. And then, actually, maybe I have some animations in here. I was not planning on talking about this at all, and I literally worked on the slides all day yesterday. So this is, ah yes, great. So they have this malicious script inside. The script will go make requests to the UK government for some reason, and then it'll download stuff from the UK government. Then, interestingly, it would take the data that it downloaded, package it into a gem, and then upload that gem again to rubygems.org.
I thought this was pretty interesting at the time, but if you read through the code that they showed on the blog post... So they said, okay, you can identify that you've been a victim of this because there will be these files on your file system. So you can find these files on your file system. And I looked at it, and when you normally install a gem, like I think we all know, when we install a C extension, for example, it will execute the extconf. So we have kind of a remote code execution problem there already, but everybody knows this.
And when I looked at their example, the malicious script was not inside of an extconf. They weren't using a C extension to do this. And so I thought to myself, like, okay, whatever. How is this impacting anybody? Even if you installed this gem, it's not going to execute itself. You would have to go into the gem and specifically do this. So this just seems weird. How does... Yes. The article never discusses how the program gets executed, and under what circumstances this program would get executed. So, I just basically said, like, "Okay, whatever. Weird."
May 16th, registration reopens. And to be honest, I just didn't think about it much anymore.
On July 6th, Luke Marshall from the Truffle Security team notified the rubygems.org team about a caching security problem on rubygems.org. And I work on the RubyGems client a bit. So I'm part of the RubyGems client team, but not the server team. But of course, we talk to each other, so we work together.
And what's interesting about this attack is... I'm going to show you the attack here. What could happen is legacy authorization was done via a GET request. So this is from the victim's perspective. So you would do a GET request to authorize the gem. So if you did gem off, it would run this. And what would happen is it would do a basic authorization. And then the response from the server would be like this, and you'd have this key at the bottom. That was your key to log in. And I want you to keep this key format in mind here for a second. Maybe a minute, more than a second probably. So that's what it looked like.
And from the attacker's perspective, it looked like this. What was happening is that Fastly was sitting in front of rubygems.org and caching the GET responses from rubygems.org. So attackers would send a request, and all of a sudden they would get a cached key as a response. So attackers don't send the authorization header, but they were able to get a valid authorization token back. So this is not good. Not good at all.
It was in production for like six years. Yeah. It's not exciting, but I don't think many people noticed it, because if you think about the behavior of this, first off, there was no evidence of abuse at the time. No evidence that this had been abused was known at the time. You have to set Accept-Encoding gzip for it to happen. And this is kind of weird because it meant that people running curl wouldn't see it, because that's not on by default. So this was just not noticed for a long time. However, Net::HTTP does turn that on by default.
And I think it also went unnoticed because, if you imagine yourself as a gem maintainer, you're going to do a gem push. So you'll do a gem push, and let's say you ended up getting somebody else's key. Like you got my key, but you're going to go upload your own gem. So you would do a gem push, and you're just going to get a weird message saying that you're not authorized to push, because you're trying to push your own gem, but you're trying to push it with my key. So it just doesn't work, and then you're like, "Oh, that's weird. Gem login." Then all of a sudden it fixes and works, and you just don't care, right? This could have happened.
So I think that's why it went unknown for so long. And also, only old RubyGems code used this, and the cache expired after 1 hour. So 3 days later they shipped a fix for this. It was completely fixed. So RubyGems shipped a fix.
September 8th, I was at work. This is my work garb. I get a DM from one of my colleagues, Emily, and she said to me, "Aaron, you're going to get an email from Gemma." Gemma is a former co-worker of mine who is working at Anthropic. And Emily said, "You're going to get an email from Gemma, and it's going to be about a really bad RubyGems hack." And I was like, you know, she could just email me, and you didn't need to DM me. And then Emily's like, "No, you're bad at email." And I was like, "Oh, right. That's true." Okay.
So I get an email from... actually, I skipped a bit of the story here. So Gemma emails me and she says, "Aaron, I'm introducing you to this person, Neve. There's been some kind of really bad RubyGems hack, and they want to talk to somebody about it." And I was like, "Yeah, sure. Okay. I mean, I can talk to you about it." And then I started freaking out, because I realized I'm not good at checking my emails. So I was like, did I miss an email? Is there something in the HackerOne that I missed? What did I do? So I'm frantically searching through my email for stuff, and there's nothing there. So I respond to the person. I'm like, I'm so, so sorry. I must have missed your email or something, but I can't find it whatsoever. This person responds to me and said, "No, there must be a mistake. This isn't a current RubyGems problem. But there's somebody, a security researcher, trying to get in touch with somebody from RubyGems, and I'm using connections to get in touch with you." And it's, I don't know, interesting. So I'm like, okay, fine.
So I get this email from a person named Sydney von Arcs, telling me that OpenAI agents had attempted to exploit this RubyGems vulnerability, the caching one. She linked to it. And it appears that these four package submissions attempted to exploit the vulnerability to steal user API keys. And I was like, this is some... Come on, you don't know. But she provided me with links, and I'm like, okay, well, you know, I'm trying to be a responsible person. I will click all of these and read them.
molite reading code. Anyway, so I go click the links and I read them, and I'm going to show you here's an excerpt from one of the gems that she sent me. This isn't the whole thing, but you go ahead and scan the QR code there that links to the whole thing, so you can go take a look at it and read it if you want to. And I'm going to point out a few things in this code. The first thing is the comment at the top that says leaked keys variants, and I start sweating in my chair, and I'm like, oh my goodness, that's not good. And you can see it's inside of this script called script.rb. And also the name of the gem is sln leaker 5. It's weird. Yeah.
Next thing I notice is that it's doing a GET request against rubygems.org. And I didn't put the path in here, but it is the authorization path. That's what that is. Then I looked at the bottom and I saw this particular regular expression, which you may have recognized from the response, the auth key
response and I realized what this is doing and I was like, oh yes, this was trying to hack rubygems.org. So I responded to them and I was like, okay, yes, that is very, very interesting. And they asked if they could have a meeting, like a meeting with me, and I was like, "Yeah, let's meet and talk about this." So, we have a meeting.
Yeah. Oh, yeah. This is me realizing what's going on here. I'm talking to them and I'm like, "How, okay, that's very interesting. You sent me this. It's again, it's in script.rb. How is this executing anywhere? Like, how would anybody execute this? It's not in the extconf." And she told me it was being run in rubydoc.info, and I was like, what? rubydoc.info?
If you go to rubydoc.info, it's just a documentation server. So it shows documentation for different gems. Let's see, what is the thing. If you look inside of the gem, you'll see that there is a .yardopts file. The .yardopts file looks like this. And you'll see that it says, like, load script.rb.
And what happens is if you have the YARD gem installed, it installs a RubyGems plugin. And then if you install another gem, when you install that gem, the YARD gem will go ahead and look inside of your gem for a .yardopts file. And if it finds a .yardopts file, it'll run this code in here. So it'll run this.
And she told me that anytime a gem got uploaded... So let's see if I have... Yes. So what was happening here is a bot would upload a gem to rubygems.org like this. rubygems.org would then send a webhook to rubydoc.info. So rubydoc.info is completely unrelated to rubygems.org. It would send a webhook to rubydoc.info. rubydoc.info would say, "Oh great, there's a new gem published." It would go download the gem.
So it would download it, run it, open it up inside of a Docker container. The Docker container still had network access. It would execute the malicious script. The malicious script would go ahead and download data from the UK government. It would package that up as a gem. It would upload the gem to rubygems.org. rubygems.org would send a webhook to rubydoc.info. rubydoc.info would execute a malicious script. It would download data from the UK government, etc., etc., etc.
So, this is crazy. Insane to me. I couldn't believe it. All of these pieces put together like this. This is wild. Like, how on earth could this be happening? I want to do kind of a timeline review on this, because this blew my mind.
So, May 11th and 12th was the gem stuffer campaign. And those gems, you can look at the upload dates on them. They are from May 11th. All right, the RubyGems cache report came in on July 6th, and the person that reported it was completely independent of OpenAI. They didn't have any affiliation whatsoever. They found it independently. So these bots knew about this problem for, I don't know, a month here, or two months, I guess. And all of this happened well before the Hugging Face breach as well.
So finally, on September 12th, Sydney and her team published rubyhack.ai, and I encourage you to read it. It has a very good in-depth look into this particular attack.
So the thing that freaked me out is that the bots figured out rubygems.org sends a webhook to rubydoc.info. They figured out that YARD will execute arbitrary code on... well, it'll execute arbitrary code. Then they figured out that rubydoc.info will process YARD docs with a network connection and execute that code.
They also figured out... this is what I suspect. You'll notice if you go look through all the gems in the gem stuffer campaign, not all of them have this code inside of them to download keys. And I think it's because when those gems started getting uploaded, folks on the rubygems.org team started shutting down the accounts and not allowing them to log in anymore. And the bots said, "Oh goodness, I need a way to log in when I don't have credentials. So, I will use these credentials." And they decided to burn a zero-day on rubygems.org.
So, this is crazy to me. They figured out this Rube Goldberg machine. Now, there's a happy ending. The rubygems.org team was able to handle it. Like, they handled it completely. Fixed it all up, closed the leaks. As far as we know, nobody was exploited, so there was no harm done to folks in the community. He's good. Yes. So, thank you so much to them.
So, I want to end this kind of here. I am an AI optimist. I use AI every single day to write code. I use it all the time. Actually, I want to change this a little bit. I'm just an optimist, not necessarily an AI optimist. I think that this is the best time to be alive, and I really believe that it's only going to get better.
But the thing is, like, I didn't have to take any pills to be an optimist. I'm actually a middle-aged man, and I don't want to take any more pills. Like, I have plenty of them. But I don't think that being optimistic about AI, or about our future, necessarily means that I need to surrender myself to it.
I've heard the argument that AI is just our next compiler. You don't need to read the machine code that your compiler produces, right? So why do you need to read the code that your AI produces? I can kind of get that argument. It's true. You probably don't read the machine code your compiler produces.
But that's because the compiler made a promise to you. It made this promise to you: the as-if rule. Whatever it does, the behavior that you can observe matches the code that you wrote. And everything that I showed you in this presentation shows you what it costs to keep that promise. Now, your AI has not made this promise to you. There is no as-if rule for your English.
So, I am not going to concede my destiny to AI. I'm going to keep reading my code, and I hope that you will, too.
And I want to end it with this slide. Get in, loser. We're going programming. Thank you.
Article published · Updated
