📊 Full opportunity report: Engineering Is Automated. Research Is the Residual. on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
AI systems are now capable of automating much of AI engineering, reaching near-saturation on key benchmarks. Research, however, remains less automated, but emerging evidence suggests it may also be increasingly handled by AI soon.
Recent evidence shows that AI systems are now capable of automating the majority of core engineering tasks in AI research, reaching near-saturation on several key benchmarks. Meanwhile, research activities, which involve more creative and exploratory work, remain less automated but are showing rapid advances, suggesting a shift in the balance of AI development efforts.
Multiple independent benchmarks, including CORE-Bench and MLE-Bench, demonstrate that AI systems have achieved 95.5% and 64.4% performance levels respectively, within 15 to 16 months. These benchmarks measure tasks such as reproducing research results and competing in Kaggle competitions, which are critical components of AI engineering. The progress indicates that automating research reproduction and engineering tasks is now feasible at scale, reducing the human effort traditionally required.
Experts note that reproducing research papers, once a significant bottleneck, is approaching a ‘solved’ engineering problem, with AI handling dependencies, code execution, and result analysis reliably. Conversely, research activities involving hypothesis generation, conceptual innovation, and long-term exploration are less mature but progressing quickly. The structural analysis suggests that research may itself be a form of large-scale engineering, implying that the residual gap could close faster than initially anticipated.
Engineering is automated.
Research is the residual.
Six skill benchmarks. Edison’s framing. The question Clark leaves open is whether research is just engineering at scale.
Jack Clark’s Import AI #455 catalogs six benchmarks measuring AI capability on AI R&D tasks and concludes “AI can today automate vast swatches, perhaps the entirety, of AI engineering.” The residual question is research. The structural read on the residual: it may not be a permanent moat.
Six skills. One trajectory.
Clark catalogs six benchmarks measuring AI capability on AI R&D-relevant tasks. Each individual benchmark could be noise. Six benchmarks moving together is a curve. The pattern is the cascade observed across the broader Clark series — visible here in the specific R&D-skill domain.

AI for Everyday Work (2026 Edition): How to Use AI for Emails, Research, Summaries & Productivity Without Technical Skills (AI Skills for the Real World Book 1)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Three data points. Mixed signal.
Clark provides three data points on the creative-spark question. Yes-evidence: Erdős-1051, centaur math discovery, sporadic Move-37-style moments. No-evidence: low yield, framing dependence, absence of acceleration. The mixed signal is the honest read.
The data supports two readings. Pessimistic: rare moments suggest creative insight is qualitatively distinct from engineering work. Optimistic: rare moments are an artifact of low-volume exploration; more shots on goal yields more discoveries. Both readings are consistent with Clark’s “vast swatches, perhaps the entirety” claim. They differ on the residual.
![Claude AI for Beginners Bible: [5 in 1] The Ultimate Guide to Automate Your Work, Save Hours Every Week, and Use AI for Real-World Results](https://m.media-amazon.com/images/I/415+fSJacsL._SL500_.jpg)
Claude AI for Beginners Bible: [5 in 1] The Ultimate Guide to Automate Your Work, Save Hours Every Week, and Use AI for Real-World Results
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Five dimensions Clark gestures at but leaves underdeveloped.
Clark’s section is rigorous on the empirical evidence. Five strategic dimensions matter for the institutional response that the Clark series synthesis argues is structurally inadequate.

AI and the Music Industry: Transforming Production, Platforms, and Practice
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Two readings. Different equilibria.
The structural question Clark leaves open: is research a permanent moat that bounds automated AI R&D, or is it engineering at scale that dissolves with more shots on goal? Both readings are consistent with the current data. They differ by orders of magnitude in consequences.
Productivity multiplier years
Recursive loop operational

The 60-Second AI Agency: How to Build and Sell Simple Automation Solutions to Local Businesses
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Five audiences. Asymmetric cost of being wrong.
The institutional response should not bet on inspiration being a permanent moat. If the distinction holds, capacity built is still useful. If it closes, capacity is necessary. Asymmetric cost-of-being-wrong points toward building now.
IN INDUSTRY
IN ACADEMIA
POLICYMAKERS
INVESTORS
EVERYONE ELSE
Engineering is automated. The residual is the question. The institutional response should not bet on inspiration being a permanent moat.
Implications for AI Development and Research Automation
This shift signifies a potential transformation in AI R&D, where engineering tasks are largely automated, freeing human researchers to focus on higher-level conceptual work. It challenges the traditional view that inspiration and creativity are the primary bottlenecks in AI progress, suggesting that institutional and strategic adjustments are needed to adapt to this new landscape. The rapid automation of research activities could accelerate innovation cycles and reshape the role of human researchers in AI development.
Recent Benchmark Progress and AI Capabilities
Over the past 15-16 months, key benchmarks such as CORE-Bench and MLE-Bench have shown exponential improvements, with AI systems now reproducing research papers at near-human reliability and achieving competitive performance in Kaggle competitions. These benchmarks serve as proxies for core engineering skills essential to AI R&D. The progress aligns with theories that AI is approaching a ‘coding singularity,’ where automation could handle most engineering tasks, leaving research as the remaining challenge.
“The evidence suggests that AI can today automate vast swaths, perhaps the entirety, of AI engineering. The residual research challenge remains, but it may be less binding than initially thought.”
— Thorsten Meyer
Uncertainties About the Future of Research Automation
While current benchmarks indicate significant progress, it remains unclear how well AI can handle the more creative and exploratory aspects of research, such as hypothesis generation and long-term innovation. The extent to which research can be fully automated, and whether inspiration remains a permanent moat, is still under debate. Additionally, the pace of institutional adaptation and the development of new benchmarks to measure research automation are ongoing.
Next Steps in Monitoring AI R&D Progress
Researchers and institutions will continue to track benchmark developments, focusing on the automation of higher-level research tasks. Expect further improvements in AI’s ability to generate novel hypotheses, design experiments, and interpret results. Industry and academia are likely to adapt their strategies to leverage increasingly autonomous AI systems, potentially accelerating the pace of innovation while reevaluating the role of human researchers.
Key Questions
How close is AI to fully automating AI research?
Current benchmarks suggest AI can automate core engineering tasks at a near-saturation level, but fully automating the creative and exploratory aspects of research remains an ongoing challenge. Progress is rapid, but complete automation is still a future goal.
What are the main barriers to automating research?
The primary barriers include the complexity of hypothesis generation, long-term strategic thinking, and the creative aspects of research, which are less amenable to straightforward automation compared to engineering tasks.
Will automation reduce the need for human researchers?
Automation is expected to shift the role of human researchers from routine tasks to higher-level strategic and conceptual work, rather than eliminate the need for human involvement entirely.
How might this shift impact AI development timelines?
If research automation accelerates as predicted, it could shorten AI development cycles significantly, leading to faster breakthroughs and more rapid deployment of advanced AI systems.
Source: ThorstenMeyerAI.com