Everyone is talking about AI. It’s noisy and political and hyped and uncertain, which makes it difficult to form your own opinion in the sea of information out there.
More recently, everyone is talking about whether it’s going to kill us all.
One of the buzzwords at the heart of this discussion is recursive self-improvement (RSI). This is the idea that once models become competent at performing research, they’ll be able to use that ability to improve themselves. AI is undoubtedly already significantly assisting human AI researchers, but RSI focuses on the ability for a model to perform research with no human assistance or intervention.
Unlike a human researcher, who needs to be born and educated and eat and sleep, an AI researcher can work around the clock and be cheaply copied to create another equally capable researcher. Taken to the extreme, this creates the possibility of a feedback loop where better models build even better models, leading to exponential growth in model capabilities.
Nobody knows if or when we’ll achieve recursive self-improvement, or how fast takeoff could be if we do. There is a wide range of possible trajectories, including one where progress looks much like it does today or halts entirely due to external shocks. This is an important question in deciding how worried to be about AI, because the speed at which models improve affects both the technological and societal benefits and risks we face and how much time we have to respond to them.
This is one of the biggest questions of our time, so it warrants far more thorough investigation than this post will provide. What I’ll do here is break RSI down into three components:
- The capabilities models need to contribute meaningfully to AI research.
- The inputs that allow those capabilities to improve.
- The feedback loop where better research produces better models, which can then do better research.
I’ll then put these components into a highly simplified model to build some intuition for how the moving parts could interact, and compare the result against the views of people who have studied the topic in much greater depth.
My aim isn’t to make a realistic prediction about whether or when we’ll see RSI, but to provide some scaffolding for building your own intuition around the topic, which can then be updated as we learn more.
Capabilities Today
We use capability evaluations to understand exactly what these models can do. These are tests designed by humans (or other models) to assess a model’s capabilities. Evaluations are imperfect, can contain errors, and don’t necessarily predict future performance. But they’re still a reasonable way to understand what models can do today.
Here are a few that I think are particularly interesting for assessing model capabilities today:
GPQA Diamond tests models on difficult biology, chemistry, and physics questions written by PhD-level experts. Human experts score around 70%, while leading models now score higher. It’s a useful signal that models can perform at an expert level on some scientific reasoning tasks.
GDPval tests models on real-world tasks across a range of jobs, like creating spreadsheets, presentations, and reports. Experts compare the model’s work to answers produced by professionals. The results are from 2025, so they’re already a little dated, but they give a useful sense of how models perform on the kinds of tasks people actually do in their jobs today.
The Epoch Capabilities Index combines several different benchmarks across a range of fields such as mathematics, reasoning and software engineering to form a single score for a model’s capabilities. This can be thought of a little like an IQ score. The number doesn’t mean much in isolation, but it reflects the trend of progress over time.
An important thing to note is that capability evaluations typically measure the kinds of problems we’re particularly good at training models to solve. If you can write an evaluation for a task, you can probably also create a learning environment where a model can practice that skill and improve. The same isn’t true for tasks with no clear measure of success, or even a clear definition of the problem. The things that make a task hard or expensive to evaluate are often the same things that make it hard to train a model to do well.
This isn’t quite Goodhart’s law. Models aren’t worse at basic arithmetic just because we can measure it. But it does mean our picture of model capabilities is incomplete. So take these evaluations as part of the evidence, but also pay attention to how models perform on some of the less defined tasks you ask them to complete.
Research Capabilities
Doing research requires more than being good at science or coding. Models need to be able to decide what questions are worth investigating, make judgement calls when there isn’t a clear answer, and coherently work on ambiguous problems over long periods of time.
Terminal bench tests models on research workflows across a range of scientific fields. At the time of writing, the best model completes around 60% of the tasks.
Anthropic has also provided some evidence of agents independently designing and running experiments within a research problem, although the overall problem and success criteria were still defined by humans. OpenAI’s recent self-reported data tells a similar story. Agent use has grown quickly for building, running, and communicating research, but much less for the harder-to-specify parts of the process, like deciding what to work on or which ideas are worth pursuing.
Another evaluation that’s relevant to the ability to perform research is task completion time horizon. Research projects run over long time horizons, so agents need to be able to coherently “self manage” to complete them.
This evaluation runs into the same issue I mentioned above: it measures well-defined software engineering tasks, so the results may not translate cleanly to long-running, open-ended research.
Capability Drivers and Bottlenecks
The scaling laws paper indicates that large language model performance increases with model size, data, and computing power. Practically, this means that we are constrained by the volume and quality of data we’re able to train on and the computing resources to run training because we’re able to design larger models if we have the resources to train them. These laws are trends extrapolated from models that we have already trained, and there is no guarantee that they will continue indefinitely.
Can we scale compute?
Our ability to scale the computing power we have available is important to RSI in two key ways:
- Training models: More compute lets us train larger models, on more data, for longer, which has historically been a significant driver of performance.
- Performing research: More compute also lets AI researchers run more experiments, simulations, and parallel attempts when trying to improve future models.
The second part may be more important than it first appears. The data is limited, but estimates suggest that final training runs account for only around 10–20% of compute spending, with most of the rest going toward research, development, experiments, and data generation. If AI researchers get much faster, they may want to run many more experiments, which means research itself could quickly become constrained by available compute.
Compute is also likely to keep growing quickly over the next few years. AI 2027 estimates that global AI compute could grow around 2.25× per year through 2027, and closer to 3.4× per year for leading labs. Epoch expects substantial scaling to remain possible through 2030, with electricity and chip production becoming more serious constraints as we get closer to that point.
AI could partly improve compute constraints by improving chip design, hardware efficiency and the operation of electricity systems. There are already examples of this happening, but it is much less clear how much AI can speed up the physical manufacturing and infrastructure buildout needed to add new compute.
Can we scale data?
There are a few different phases involved in training a model, each contributing to what the model can do in different ways:
- Pretraining: Train the model on massive amounts of text, usually the most of the internet. This gets the model to the stage where they are an incredibly smart autocorrect that can predict the next word in a sentence.This stage is where the model builds most of its knowledge and capabilities.
- Mid-training: Continue training the model on more curated data, like scientific textbooks, code, or worked math problems, to improve its capabilities in particular areas. This uses a similar process to pretraining and helps to build specific knowledge and skills.
- Supervised fine-tuning: Train the model on examples of the kinds of responses we want, like conversations between a user and an assistant. This teaches the model how to use its existing knowledge and skills in a particular setting.
- Reinforcement learning: Let the model try a task many times, then tell it which attempts were better or worse. Over time, it learns which approaches are more likely to lead to a good outcome. This can make existing capabilities more reliable, or teach the model to use what it already knows in new ways.
The scaling laws referenced above mostly apply to pretraining, and likely carry over to mid-training because it uses a similar process. We don’t yet have the same clear, general rules for how the other processes scale into better capabilities.
It’s estimated that we could run out of high-quality human-generated text for pretraining by around 2030. We could use existing models to generate “synthetic data”, but this only helps if it gives the model something new to learn. This is more straightforward in defined areas like mathematics, where models can generate new problems and check their answers, and much less certain in messier domains where there is no clear way to tell whether the generated data is actually good.
The data we can capture also needs to contain the things models need to become good researchers. Papers and textbooks contain a huge amount of scientific knowledge, but research also relies on things that are harder to write down: choosing which questions are worth pursuing, noticing when a result looks strange, or knowing when to abandon an approach. We mostly record the outputs of successful research, not all of the judgement that produced them.
Models may still be able to learn some of this indirectly. They can combine ideas and generalize from what they’ve learned elsewhere, so they don’t necessarily need to see every skill demonstrated directly in their training data. The uncertainty is how far this goes.
We could also try to teach these harder-to-capture parts of research through reinforcement learning. Models can be put in environments where they propose ideas, run experiments, and learn from the results. The difficulty is that this works best when we can define what success looks like, and good research is often hard to score.
A note on Algorithms
(A warning for ML researchers: what follows is intentionally simplified)
Improvements to the efficiency of algorithms we use to train models can decrease the amount of compute needed to reach the same level of performance, which means we can train more capable models on the same hardware. It’s difficult to decouple these gains from improvements in data and increased scale because all of these tend to change together. Available estimates have a very large margin for error, but suggest that improvements to AI software may be equivalent to several-fold more effective compute each year.
Right now, frontier models are built using transformers. They were a significant step forward in performance, and were developed by trying different ideas and seeing what worked, guided by theory and intuition. There’s no reason to assume that transformers are the best architecture we could ever find; they’re just the best option we know of.
AI-driven research could run experiments, measure results and improve algorithms in the same way people have. Models are likely a long way from independently discovering a completely new architecture, but larger improvements seem plausible if they’re given enough time and compute to experiment.
The Feedback Loop
What makes RSI different from ordinary research is the potential for a feedback loop. A human researcher who makes a breakthrough moves their field forward, but they’re still the same researcher afterwards. An AI that makes a breakthrough can produce a better model, and that better model may be able to make the next breakthrough faster.
If each model speeds up the next, then you get a takeoff. If something slows this down, then you get something closer to today’s progress. Parts of this loop can be measured, but the whole thing can’t yet. So this section leans more on educated guesses from researchers, and on what the labs say about themselves, than the rest of the post.
Research Boost
The loop only gets going if better models actually help with the work of building their successors. Each new model can make the people and agents working on the next one more productive.
The labs report that this boost is already large, although all note that these measurements are difficult to isolate. Anthropic says more than 80% of its merged code is now written by Claude, and its researchers report about 4× more output, up from +50% a year earlier. OpenAI reports 3.1 agent-workdays for every human workday.
More code and more experiments aren’t the same as faster progress, but even a modest boost compounds if every new model makes the next round of research a little faster.
Research Autonomy
How much of the research AI can take on without people matters as much as how fast it works. Amdahl’s law says a process can only go as fast as its slowest part. If AI does 95% of the research but humans still do the other 5%, research can only go about 20× faster, however good the AI gets.
In OpenAI’s breakdown, agents mostly build and run experiments, while humans still choose priorities and decide whether to scale or ship. METR’s timelines model puts automation at 25–50% of AI research tasks in early 2026, with a median of more than 99% by late 2032. The last few percent matter a lot: going from 90% to 99% moves the ceiling from 10× to 100×.
I think that this will come down to research taste and judgement. If AI can make good decisions about what to research, the loop can run much faster. If humans are still needed for those decisions, it stays tied to human speed.
Deployment Lag
A good idea doesn’t help the loop until it ends up in a model that can use it for the next round of research, and every turn of the loop waits on that journey.
OpenAI describes this development cycle as including experiment design, evaluation, infrastructure work and debugging before changes make it into training. Epoch estimates frontier training runs already take on the order of several months. RSI could compress much of the work around the training run by automating experiments, coding, evaluation and debugging, but some latency remains because experiments and training still have to physically run.
Deployment lag therefore acts less like a ceiling on takeoff than a brake on its timing. It doesn’t necessarily change how far capabilities eventually rise; it spreads those gains over more calendar time, turning what might otherwise look like a near-vertical jump into a sequence of fast steps over time.
Hardware Speedup
AI could also accelerate the loop by improving the hardware that makes training and research possible. That could mean designing better chips, getting more useful compute out of existing hardware, or making data centers more efficient to build and run so that the compute bottleneck is eased.
AI is already being used to improve chip design, but turning those gains into substantially more compute still requires physical infrastructure. Epoch estimates roughly 1–2 years to build a data center, 2–3 years for very large facilities or power plants, and 4–5 years for a new cutting-edge chip fab.
RSI could shorten design and planning, and sufficiently capable robotics could eventually speed up construction and manufacturing too, but for now hardware looks like a much slower feedback loop than software.
Difficulty of Innovation
Not everything pushes the loop faster. As the obvious ideas get used up, each improvement may take more research than the last.
We don’t have good evidence for how quickly this effect sets in for AI research. Moore’s law is probably the closest analogue we have: sustaining progress in chip performance has required far more researchers over time. But AI research may behave very differently, so this is at best a rough comparison rather than something we should assume carries over.
This is one of the main things that determines whether RSI runs away or simply makes progress much faster. If better AI researchers increase research productivity faster than new ideas become harder to find, the loop accelerates; if not, automation can still compress years of progress without producing an explosion.
Back of Envelope
To build some intuition for how these pieces interact, I’ll bring them together in a highly simplified model. This is not intended as a prediction, but as a way to understand the relationships between the moving parts and update your intuition as new information comes in.
To interact with the model, you can choose a reasonable starting state based on where you think we are today, and set your beliefs about the capabilities and bottlenecks we’ll see in the future.
Expert Projections
It’s difficult to form your own opinion here. The types of problems you use AI for will bias your view of how it’s progressing, and investigating the many moving parts in RSI requires more time than most people have. Expert opinions can therefore be useful as another input, helping to anchor your intuition against people who have spent much more time studying the question.
These opinions aren’t ground truth either. People have different models of how progress works, different access to frontier systems, and their own incentives and biases. I think the useful approach is to look across a range of informed views, understand why they differ, and combine them with your own reasoning rather than relying on any single forecast.
The chart below gives a sample of those views, ranging from relatively skeptical to much faster takeoff scenarios.
Conclusion
Predictions on AI timelines tend to age like milk, so I expect to update these views as my understanding grows and new information comes in. My current intuition is that AI will become extremely useful for performing and speeding up research, but that data, compute and the messy realities of doing good research will create bottlenecks along the way. I’m not expecting an exponential takeoff. But I am expecting progress to be very, very fast.
Which means there is work to be done. Throughout this post, the evidence that matters most has been the hardest to get: evaluations measure what’s easy to measure, the best numbers on research uplift come from the labs’ own reports, and nobody has measured the feedback loop as a whole. We need better ways to measure the messy, open-ended capabilities that research depends on, and we need labs to publish evidence about their progress that others can check. Most of all, the loop gets faster precisely as humans step out of it, which is also when it’s hardest to notice if models are taking shortcuts rather than doing what we actually want.
Many thanks to Alexander Reinthal for helping me out with helpful resources, questions and review. Also thanks to Elle Mouton and Leila Stein for their input.