TL;DR
A claim that an AI system scored 44% on the ARC-AGI-1 reasoning benchmark using roughly 67 cents of inference compute is attracting wide attention. The benchmark itself and its history of cheap high scores are well established; the specific result and who achieved it remain unconfirmed.
A claim circulating online that an AI system scored 44% on the ARC-AGI-1 benchmark while spending only 67 cents in inference compute has drawn a spike of attention across AI research communities and social platforms. The benchmark at the center of the claim is real and well documented, but the specific result — who produced it, with what model, and under what verification conditions — remains unconfirmed at the time of writing.
The phrase “44% on ARC-AGI-1 in 67 cents” refers to a reported score on the Abstraction and Reasoning Corpus, a puzzle benchmark created by François Chollet to test fluid reasoning — the ability to generalize from a handful of examples to novel problems, rather than pattern-matching on memorized data. ARC-AGI-1 was run for years as a public competition by the ARC Prize Foundation, and it has long served as a shorthand test of whether a system can do something closer to human-like abstraction than standard machine learning benchmarks capture.
What is established fact: cheap, high-performing entries are not new to this benchmark. In late 2024, DeepSeek’s open-weights reasoning models produced ARC-AGI-1 results at a small fraction of the compute cost of frontier proprietary systems, a result widely covered and credited with reshaping assumptions about how much money is needed for strong benchmark performance. The idea that a strong ARC-AGI-1 score can be obtained for under a dollar in inference is therefore plausible given the benchmark’s history, and is likely one reason the current claim has spread quickly.
What is not established: the identity of the system behind the current “44% in 67 cents” figure, the exact experimental setup, whether the score was achieved on the benchmark’s semi-private evaluation set (the standard for verified claims) or on the public training tasks, and whether the cost figure reflects the full pipeline including any test-time search or repeated attempts. No verified leaderboard entry, preprint, or institutional announcement tied specifically to this figure has been identified as the source of the current wave of interest.
Why Cheap Benchmark Scores Rattle the Field
If verified, a 44% ARC-AGI-1 score at 67 cents of compute would land in a meaningful performance band on a benchmark deliberately designed to resist brute force. ARC-AGI-1’s public leaderboard showed for years that most machine learning approaches plateaued well below human-level performance, which is around 85% for typical human solvers. A cheap 44% would sit comfortably above many far more expensive historical attempts, reinforcing the argument — already strengthened by DeepSeek’s 2024 results — that cost-efficiency, not just raw capability, is where AI progress is concentrating.
The claim also matters because ARC-AGI-1 has become a proxy argument in debates over benchmark saturation. As mainstream benchmarks fall to top models, ARC-style tasks are held up as evidence that general abstraction is still unsolved. A low-cost near-mid score complicates that narrative: it suggests the remaining gap may be as much about economics and engineering as about fundamental capability limits. That framing is interpretation, however, and depends entirely on the claim holding up under verification.

Distributed AI Systems: A practical guide to building scalable training, inference, and serving systems for production AI
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
ARC-AGI’s History of Upsetting Cost Assumptions
ARC-AGI-1 was introduced by François Chollet in 2019 as part of his paper “On the Measure of Intelligence.” The benchmark presents visual grid puzzles that require inferring an underlying transformation rule from a few examples. Unlike ImageNet-style benchmarks, it was explicitly constructed so that memorizing training data provides little benefit.
The benchmark gained mainstream visibility through the ARC Prize competition, which ran from 2020 and offered significant prize money for verified progress. Early winners relied on hand-engineered program synthesis rather than neural networks. The landscape shifted in late 2024, when reasoning models — most prominently from Chinese lab DeepSeek — achieved high scores at inference costs orders of magnitude below what US frontier labs reportedly spent, a development widely discussed as a challenge to the assumption that top-tier reasoning requires top-tier budgets. ARC-AGI-1 has since been succeeded by ARC-AGI-2, a harder successor benchmark, which is where most current verified competition activity now takes place.
machine learning inference cost optimizer
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the 67-Cent Claim Leaves Unverified
Several things about the current claim are unknown. First, the source is unidentified: no verified leaderboard submission, paper, or vendor announcement has been confirmed as the origin of the figure. Second, the evaluation conditions are unspecified — whether the score was measured on the semi-private evaluation set that the ARC Prize Foundation uses for verified results, or on public tasks where overfitting is possible, materially changes how the number should be read. Third, the cost accounting is unclear: “67 cents” could mean a single-pass inference cost, an average over a test suite, or a cost that excludes test-time compute such as sampling multiple candidate answers. Finally, it is not confirmed whether the system involved is a known model, a fine-tune, or a novel approach. Until these details surface, the figure should be treated as an unverified claim, not a result.
As an affiliate, we earn on qualifying purchases.
How This Gets Confirmed or Discarded
The most likely path to verification is an entry on the official ARC Prize leaderboard or a technical report detailing the method and cost accounting. The ARC Prize Foundation has historically required semi-private evaluation for prize-eligible results, so a legitimate claim would eventually surface there. If the figure came from a research team, a preprint or blog post with reproduction instructions would be the expected follow-up. Readers tracking this should watch for a named system, a documented evaluation protocol, and independent replication — the three elements that separate a real result from a viral number.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is ARC-AGI-1?
It is a benchmark of visual reasoning puzzles created by François Chollet in 2019, designed to test whether a system can infer abstract rules from a few examples and apply them to new problems — something humans do easily but that resists memorization-based machine learning.
Is the 44% score for 67 cents confirmed?
No. The claim is circulating widely, but no verified leaderboard entry, paper, or institutional announcement tied to this specific figure has been identified. The system, evaluation method, and cost accounting are all unconfirmed.
Is 44% a good score on ARC-AGI-1?
It is a solid score by historical standards. Top reasoning models have exceeded it, and typical human solvers score around 85%, but many far more expensive systems historically scored lower, which is why the cost figure draws attention.
Has a cheap high score on ARC-AGI-1 happened before?
Yes. In late 2024, DeepSeek’s open reasoning models achieved strong ARC-AGI-1 results at inference costs far below those of frontier proprietary systems, an established and widely covered precedent for cost-efficient benchmark performance.
How would the claim be verified?
Through an official ARC Prize Foundation leaderboard submission using the semi-private evaluation set, or a published technical report with enough detail — named model, protocol, and cost breakdown — for independent replication.
Source: hn