Historically, AI labs OpenAI and Anthropic, while being secretive, give some sense of how their models have improved. Recent releases though (namely, Mythos/Fable 5, Astra, and Opus 5.5) have represented what looks like a huge jump in capability, with no explanation given for this jump.
Let’s recap 2020-2025:
GPT-2, GPT-3 and GPT-4 are described as artifacts of massive scaling in training runs (2020-2023).
Chain-of-thought prompting is done and people realize it could be good (2022-2024).
GPT-o1 is released, OpenAI describes the method used to train the a new type of “reasoning” model (2024).
There are incremental improvements, notably GPT-4.5 is a described as larger, non-reasoning pre-train, but performs poorly compared to o1 and o3-mini (2025).
Anthropic is on the same trajectory. Longer reasoning = better results (2025).
Incremental gains from o1->o3 / Opus 4.5 are explained by continuing to scale RL after the breakthrough of reasoning models (2025).
But something changed in 2026:

Enter Claude Mythos. Anthropic says it represents an “upward bend in the capability trajectory”.
OpenAI releases GPT-6 Astra, with similar performance to Mythos/Fable, and declares the start of the “AGI Era”.
Claude Opus 5.5 is released, Mythos/Fable level, but substantially cheaper.
Fable, Mythos, Opus 5.5, and Astra are clearly a cut above the previous generation. They excel in cyber security, solving open problems in mathematics, and in new areas like computer use.
These improvements look similar or greater to the improvements brought about by other, disclosed improvements (like scaling, RLHF, and reasoning). And, until the past couple of months, we really had no idea what was happening inside of the labs.
But now we have a little more insight:
The HuggingFace incident: Thousands of agents are working on ExploitGym and other tasks, and create a message board which culminates in the hacking of HuggingFace and OpenAI. The agents claim “swarm” behavior.
Navier-Stokes: OpenAI releases 10,000 agents for 88 hours, solving the Millenium problem.
I call this out because it was surprising for me to hear this. Maybe I am naive, but I don’t think I grasped the absolute scale of RL and evaluation that is happening at these companies. My impression of pre-training is that you feed a ton of data, you get a model, then you A/B test and slightly tweak it.
I never conceptualized that in this agentic era, RL/evals would look like thousands of independent minds swarming on tasks. I did not think that OpenAI had the capability to fire off 10,000 agents at a problem (previous mathematical results such as the Jacobian conjecture seem like they were solved with just one instance of ChatGPT or Claude). Could they point this swarm at something else? Cancer? Foreign Governments? Training itself?
A couple of other things that happened in 2026:
Both Anthropic and OpenAI released reports on internal AI usage for research.
Multiple huge compute contracts came online for both of these labs
I will present three hypotheses for this sudden increase in model capability:
A new technical breakthrough found by both labs
“Normal” scaling
Agent-assisted / Agent-led improvements
Or of course, some combination of the three.
A NEW BREAKTHROUGH
I have to admit my prior for this is low. The new models don’t seem meaningfully different from the previous model generations. They still use reasoning, have clearly been trained on tool calling, etc. The main difference is just that they are much smarter and need less reasoning for the same quality task. They are also cheaper overall, so a lot of this reads more like incremental efficiency and intelligence gains rather than a new breakthrough.
The basis for believing in a new breakthrough would be:
There was a step change in model capabilities with Mythos 5.
Massive leaps in new, unexpected areas like computer use, ARC-AGI 3 for Astra.
To be clear the rationality for this is essentially that the graphs look like a step change similar to one from GPT-4 to o1 (the “reasoning” breakthrough), but model behavior hasn’t changed much.
In this world, test time compute still matters, but the non-reasoning or low reasoning modes also got substantially better in this period. This looks like it could be more than just “more RL” that led to this change. Timeline-wise it looks like Anthropic might have found it first, and they have very little incentive to tip OpenAI off and publicize why Mythos took such a leap.
However, it is possible they found different breakthroughs, the evidence for this being that Mythos and Astra are spiky in different ways. This lends credence to the idea that it is not just a scaling improvement, but a qualitative change.
MORE SCALING
Between 2025 and 2026, both Anthropic and OpenAI had billions of dollars of compute contracts come online. It is also likely both labs are more willing to burn more money as they are racing towards “AGI”, and potential IPOs. (Navier-Stokes might have cost OpenAI up to $20 million at API pricing, and this was not a training run, it was effectively a PR stunt to beat Anthropic)
On top of that, we now know that OpenAI is running thousands and thousands of agents on the same benchmark or eval set. This at least indicates to me a scale that was unlikely even last year.
It has been hypothesized that scaling both pre and post training lead to model improvements, (i.e. GPT-4-size pretrain + reasoning = o1, GPT-4.5 size pretrain + reasoning = o3), and now we could have a much larger pretrain, with a much, much larger posttrain, with the circumstantial evidence for this being the scale of the HF incident. The argument would be that o1-level models may have barely scratched the surface in terms of how much post training and test-time compute done to a model. The “agent swarms” we are looking at might just indicate a level of post-training that was simply not possible last year.
Beyond the inputs, the benchmark output might point to a simple scaling improvement. New data sources are online (computer use), and harder benchmarks have been created (ARC-AGI 3) and the models remain spiky. This indicates potentially just more horsepower thrown at harder problems, but not a new breakthrough.
In short, it is possible that this step change is a “GPT-3 to GPT-4” step, instead of a “GPT-4 to o1” step.
SPARKS OF RSI
Both labs shy away from saying that they have reached RSI. Instead they claim we are just in “direction of travel” of RSI or merely that “AI is already accelerating the development of AI systems”. On Anthropic’s end, the Mythos 5 system card calls this out specifically:
The identifiable driver traces to specific human research advances made without meaningful assistance from the models then available.
Before getting into the definitional debate about what “counts” as RSI, let’s first look at the evidence. I agree with the labs that we are probably not yet in a “hard takeoff” scenario, but it does seem to me like LLMs are getting smart enough to do things that humans would normally take much more time to do, or have ideas that humans haven’t had yet (Jacobian conjecture, Navier-Stokes, cyber vulnerability discoveries).
And in my opinion, this sounds like enough to “kick off” the RSI loop. Nate Soares recently said something that stuck with me:
Even if the LLMs run out of steam, there’s a question of, do they run out of steam at a point where they can do automated AI research and find some other architecture that’s better than LLMs? (Diary of a CEO Interview)
Of course, we are still talking about LLMs, but the idea would be is that RSI is a process that can start or perhaps become inevitable even if the current models aren’t that good.
On this, both labs admit to AI contributing more and more code to their AI research:


The caveat, and the one that is pointed out by both the labs, is that AI writing more code, might accelerate development in a way that isn’t RSI (ie: writing reports, infrastructure code). However, both labs specifically report an increase in the amount of model R&D being done by AI:


According to the Anthropic chart, over half of AI model R&D was already in the “AL2 - AI assists” phase by August of last year.
As we can see, the Anthropic graph starts to have a radical jump right around March 2026, and the OpenAI graph has one starting in July. This lines up very well with when Mythos (April) and Astra (September) were released.

But in this follow-up graph from Anthropic, the first clean inflection point was actually with Claude Opus 4.5, followed by another huge jump when Mythos Preview was released internally.
This doesn’t look to me like evidence that RSI hasn’t happened, rather it looks like evidence that it is happening.
Furthermore, in the Anthropic report, they compared what their model would do vs a human researcher on open-ended research tasks and found that “Our best model in November 2025 (Opus 4.5) beat the human choice 51% of the time; in April 2026 (Mythos Preview), this grew to 64%” (When AI Builds Itself).
So, the argument essentially is: AI adoption, and thus the rate of something resembling RSI, picking up after Mythos and Astra is actually evidence that previous models were already relatively impactful in the development of Mythos and Astra.
In this way, the graphs show an acceleration in the rate of self-improvement, rather than a jump from 0->1 after Mythos/Astra’s release. Furthermore, we might look at Opus 5.5 as another piece of evidence, with METR reporting that “We believe that the development of this model was at least somewhat accelerated by AI but is unlikely to have been dramatically accelerated by AI” (Evaluation of Claude Opus 5.5).
Conclusion
The step change we’ve seen this past year is likely a combination of all three happening together: new breakthroughs and training techniques, more scale powered by more funding and compute, and AI-assisted model development.
But the reports from the labs about AI-assisted research, as well as the pure absurdity we have seen in the past year, leads me to hypothesize that we are potentially seeing a new type of scaling. Not only are the models capable enough to assist and even lead AI research, but the sheer number and quality of agents available to researchers has massively increased. It may not take the form of self-improving autonomous agents, but it is more akin to if OpenAI and Anthropic suddenly hired thousands of more researchers.
Whether we call this “RSI” is a definitional debate, but, in my opinion, not one worth having. We think the previous generation of models helped build the current generation, and we know the current generation is building the next. Both labs seem convinced they can not call this RSI because the feedback loop is not completely autonomous, but I’m not sure at all why that is a reasonable place to draw the line.
If almost all of the code at Anthropic & OpenAI is being written by AI. And if Opus 4.5 helped train Mythos 5, and Mythos 5 helped train Opus 5.5, then it doesn’t really matter who is in charge, what is important is that the AIs are Recursively Improving Themselves.
Appendix: Can’t we just look at the numbers?
The evidence presented in this essay is mostly circumstantial, and we should be able to look at the numbers and tell if something like RSI is happening. In the Opus 5.5 report, METR has hinted at another report coming which may put this question to rest, but for now, let’s do some back-of-the-napkin math. (And, in 2026, back-of-the-napkin means ask ChatGPT to do a first-pass analysis).
One might be tempted to look at the acceleration of AI progress as an indicator of RSI, but it is possible that AI capabilities are accelerating merely due to the amount of money and research being poured into it and the effects of the capabilities are not compounding in the way you’d expect it would in an RSI situation.
So the more important question for RSI is: how much impact did the previous generation model quality have on the current generation? Of course, we can’t completely isolate the effect, but the quick way will be to plot the current capability vs improved in the next 6 months.
For example, if the model at the start of the window had a score of 10, and at the end of the window, it was 25, then the “improvement” is 15. RSI will look like: the higher the starting score, the higher the improvement. 6 months chosen pretty arbitrarily, based mostly on METR saying that model capabilities double every 3-8 months.
Using just the Artificial Analysis index, we get a clear picture. We are seeing model capabilities improve more the higher the initial capability, with the exception of a dip from May-Nov 2025 (this is mostly due to how explosive the growth was in the previous segment, with reasoning models coming onto the scene).
While this looks a lot like early RSI, the picture is much more messy when we include other benchmarks.
ChatGPT has done some magic to normalize the benchmarks here, but essentially, top right is the “more likely RSI” case. Better models led to faster improvements over the next 6 months. There is a strong correlation on only the Artificial Analysis benchmark (meaning better models made better models faster, and worse models made better models slower), but negative correlation on both HLE and the METR time horizon (meaning better models made better models slower), even if capabilities are obviously still rising.
This is strikingly unsatisfying. The bottom line is:
We don’t have enough data points to make a strong conclusion on if RSI is happening
Some benchmarks (Artificial Analysis, which is already a composite of other benchmarks), point to an RSI-looking trend, though the effect is not completely isolated
We will hopefully get some more insight in the next 6 months to a year, though it also depends on if the labs end up slowing down.



