📊 Full opportunity report: Every Benchmark Launched 2023-2024 Has Fallen — The METR / SWE-Bench / CORE-Bench / MLE-Bench / PostTrainBench Sequence on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Six key AI benchmarks launched between 2023 and 2024 have all reached saturation or are close to it within months. This pattern suggests a rapid acceleration in AI capabilities, impacting forecasts and industry expectations.

All six major AI capability benchmarks launched in 2023 and 2024 have either been saturated or are approaching saturation within a few months, marking an acceleration in AI research capabilities.

According to Thorsten Meyer and Jack Clark, each of these benchmarks—measuring different aspects of AI research and engineering—has reached or is nearing saturation on a timeline of months rather than years. For example, the SWE-Bench, which assesses real-world software engineering tasks, improved from 2% to 93.9% in 30 months. Similarly, the METR time horizon benchmark, measuring task duration, expanded from 30 seconds to 12 hours over four years, a 1,440-fold increase. The CORE-Bench, evaluating research reproduction, was declared solved in December 2025 after improving from 21.5% to 95.5% in 15 months. These patterns have been consistent across all six benchmarks, with each showing rapid improvement and nearing saturation.

Experts note that this pattern indicates a notable shift in AI development, with capabilities advancing quickly. The saturation of these benchmarks suggests that AI systems are making substantial progress in core research and engineering skills, raising questions about the pace of future progress and the potential for AI to reach or surpass human-level performance in various domains.

Impact of Benchmark Saturation on AI Development Trajectory

The rapid saturation of these benchmarks indicates that AI research capabilities are progressing quickly. This may influence forecasts predicting AI reaching certain milestones by 2028, including autonomous research and development. For industry, policymakers, and researchers, it suggests that AI systems could soon perform complex research tasks that previously required human expertise, potentially affecting innovation cycles, labor markets, and technological competition. However, it also highlights the need to evaluate whether current benchmarks adequately measure the full range of AI capabilities.

The Senior Engineer’s AI Agent Reference: 40 Production Architectures with Failure Modes, Cost Benchmarks, and Observability Runbooks

The Senior Engineer’s AI Agent Reference: 40 Production Architectures with Failure Modes, Cost Benchmarks, and Observability Runbooks

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Benchmark Development and Progress

Since 2022, multiple benchmarks have been introduced to measure AI research and engineering skills, with the goal of tracking progress toward autonomous AI capabilities. These benchmarks were designed to be challenging and to reflect real-world tasks, including software engineering, research reproduction, and AI fine-tuning. Over the past three years, each benchmark has shown rapid improvement, culminating in saturation or near-saturation in 2025 and 2026. Notably, the SWE-Bench improved from 2% to nearly 94%, and the CORE-Bench was declared solved in late 2025. The pattern across these benchmarks suggests that AI systems are quickly closing gaps in core capabilities, driven by advances in model architectures, compute, and training techniques.

“Every benchmark launched in 2023-2024 has saturated or is nearing saturation within months, indicating a structural acceleration in AI research capabilities.”

— Thorsten Meyer

AI Engineering: Building Applications with Foundation Models

AI Engineering: Building Applications with Foundation Models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Uncertainties About Benchmark Validity and Future Capabilities

It remains uncertain whether these benchmarks fully capture the capabilities of emerging AI systems or if saturation indicates a plateau. Critics suggest that benchmarks may be overfitted or that AI systems are exploiting loopholes, which could affect the interpretation of progress. Additionally, the long-term trajectory beyond 2026 is still uncertain, with questions about whether further improvements will follow the same rapid pattern or slow down as systems approach theoretical limits.

Sunnytech Mini Hot Air Stirling Engine Motor Model Educational Toy Kits Electricity HA001

Sunnytech Mini Hot Air Stirling Engine Motor Model Educational Toy Kits Electricity HA001

Well Made—-Produced with high quality material and adopts mirror surface stainless steel for the base, brass cylinder premium…

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Monitoring AI Progress and Benchmark Evolution

Researchers and industry analysts will continue to monitor the saturation status of these benchmarks and develop new, more challenging tests to measure AI capabilities. Further validation of whether these benchmarks accurately reflect real-world performance is expected. Additionally, policy discussions are likely to focus on the implications of rapid AI advancement, including regulation, safety, and ethical considerations, as the pace of progress continues.

XTOOL D9S 2.0 ECU Coding Bidirectional Scan Tool, AI-Assisted OE Level Automotive Scanner, Topology Map, PMI, 45+ Resets, FCA/CAN FD/DoIP, Professional Full System Diagnostic Scanner, 3 Years Update

XTOOL D9S 2.0 ECU Coding Bidirectional Scan Tool, AI-Assisted OE Level Automotive Scanner, Topology Map, PMI, 45+ Resets, FCA/CAN FD/DoIP, Professional Full System Diagnostic Scanner, 3 Years Update

Empower OE-Level Automotive Scanner with 3-Yr Free Updates – XTOOL D9S V2.0 automotive diagnostic tool: XTOOL Scan Tool…

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does benchmark saturation mean for AI development?

It indicates that AI systems are reaching or surpassing the capabilities these benchmarks measure, suggesting rapid progress and potential achievement of human-level performance in specific tasks.

Are these benchmarks reliable indicators of overall AI progress?

While they are designed to be challenging, there is ongoing debate about whether benchmarks fully capture AI capabilities or if saturation reflects overfitting or loopholes exploited by models.

What are the implications of these findings for the AI industry?

The rapid saturation suggests that AI systems may soon be capable of autonomous research and development, which could influence innovation cycles, but also raises safety and ethical considerations.

Will new benchmarks be introduced to replace saturated ones?

Yes, researchers are expected to develop more complex and comprehensive benchmarks to continue measuring AI progress beyond current saturation points.

When might we see AI systems surpass human-level performance in research tasks?

Based on current trends, some forecasts suggest significant milestones could be reached by 2028, but uncertainties remain about long-term capabilities and evaluation methods.

Source: ThorstenMeyerAI.com

You May Also Like

SpaceX Owns Every Layer of AI Now. The Model Is Still the Weak Link.

SpaceX has bought Cursor for $60 billion, gaining control of all AI layers but still relies on a weak model. The move consolidates industry power.

GPT-5.6

OpenAI has officially launched GPT-5.6, featuring improved safety measures and performance updates, marking a significant step in AI development.

ShinyHunters · The New APT Model.

ShinyHunters has evolved into a new operational threat actor, combining AI-enabled tactics, a collective structure, and scalable monetization, marking a shift from traditional APTs.

The Menu: What Ten Answers Reveal

An analysis of ten jurisdictions’ responses to automation, highlighting key differences in income, capital, work, skills, and institutions, and what they reveal about policy options.