AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Livenerf, an independent benchmark tracking Claude Opus 5.5 for performance changes, has collected six of 30 planned daily runs and has not reported a post-launch decline. Its first possible comparison with the launch-week baseline is expected around October 24, 2026.

Livenerf has not yet found whether Claude Opus 5.5 was nerfed: the independent benchmark had collected six of 30 planned daily runs as of September 29, and its first comparison with launch-week performance is not expected until around October 24. The project is designed to test claims that a model becomes less capable after release, but it has published no measured post-launch change so far.

The benchmark began collecting data on September 24, about two and a half days after Opus 5.5 launched on September 22, according to the project’s GitHub report. Livenerf plans to run once a day for 30 days: the first 10 days establish a baseline, and the following two 10-day periods are compared against it. The first results row will appear after day 20, meaning an initial comparison may be possible around October 24.

As of September 29, six of the 10 baseline days had been collected, with no days missed. Each run completed 90 samples using the same benchmark harness and pinned Claude Code CLI version, the project reported. Day five required one override of a budget guard, which is recorded in the project’s deviations log.

The test uses a pre-registered panel of 78 questions drawn from GPQA Diamond, MMLU-Pro, competition mathematics and AIME 2025–26. Livenerf says it selected questions that Opus 5.5 sometimes answers correctly and sometimes misses, then measured how selection affected the fresh-sample pass rate. The protocol tracks accuracy against the baseline and median output tokens, which may shift if a model spends less effort even before accuracy changes.

At a glance
updateWhen: Status reported September 29, 2026; fir…
The developmentAn independent benchmark tracking Claude Opus 5.5 has reached six of 30 planned daily runs, but has not yet produced a result showing whether the model’s performance changed after launch.

How the Benchmark Can Detect Drift

Livenerf addresses a recurring dispute about whether a model’s quality changes after release. Users may perceive a decline, but without measurements taken near launch, later comparisons can be hard to interpret. This project’s day-one collection gives it a baseline for the specific Opus 5.5 version it is following.

The benchmark’s reach is limited. Its validation suggests it can detect an accuracy change of about 7.5 percentage points over a 10-day window, at a reported cost of about 3.6% of the weekly plan. That sensitivity means smaller shifts could go undetected. The project also says its validation could not distinguish Opus 5 from Opus 5.5 at 99% confidence; it has not shown that a 10-day window would detect a swap of that size.

Any eventual result will apply to this test setup and question panel. It would provide evidence about measured performance on those items, not establish why a change occurred or prove a particular change to model weights, routing or serving.

Amazon

Top picks for "livenerf opus nerf"

As an affiliate, we earn on qualifying purchases.

A Baseline for Opus 5.5

Livenerf’s stated aim is to test reports that Anthropic models may perform worse days or weeks after launch. The project lists possible explanations—including quantization, routing changes, lower effort or a smaller model behind the same name—but presents them as possibilities, not established causes. It also allows that perceived changes could reflect noise.

The project says ordinary sampling settings are unavailable in its current setup and that model responses cannot be made fully deterministic. It instead holds other parts of the process steady with frozen prompts, a pinned CLI, fixed graders and retained raw logs, then uses statistical comparisons. The benchmark runs through Claude Code using a Claude Max subscription and is built on the Inspect evaluation framework.

Livenerf reports that its initial screening covered 2,336 questions, with four samples per question. It says Opus 5.5 answered about 93% correctly on the first try, while 97% of questions were consistently right or consistently wrong. The 78 questions with mixed outcomes make up the calibrated panel. A report-only audit flagged possible answer-key errors and ambiguities, which the project says it has retained for a pre-registered sensitivity analysis rather than removing from the benchmark.

What the First Runs Cannot Show

No post-launch score comparison is available yet, because the benchmark is still collecting its baseline. The six completed days show that data collection has proceeded, not whether Opus 5.5 has improved, declined or remained stable since launch.

The project’s detection limits also matter. It estimates that accuracy shifts smaller than roughly 7.5 points over a 10-day window may be difficult to identify. Output-token changes may offer an earlier signal of altered effort, but token counts alone would not establish a change in accuracy or explain its cause.

Livenerf says its safety classifier sometimes routes responses through Opus 5 or refuses certain biology and mathematics questions; affected samples are rejected and counted, and questions the classifier touched are excluded. The available report does not establish whether serving behavior outside this benchmark has changed, or whether any future score movement would reflect the model, the serving path, or measurement noise.

First Comparison Due in October

Livenerf plans to continue its daily runs until it has completed 30 days of collection. After day 20, it expects to publish the first comparison of a 10-day window against the launch-week baseline, with later windows adding more observations. The report says the results table will show negative and positive changes alike, alongside uncertainty estimates and output-token data.

The next useful milestone is therefore the first comparison, expected around October 24 if the schedule holds. Until then, the project’s status is a running benchmark with six baseline days recorded—not a finding that Opus 5.5 has or has not been nerfed.

Key Questions

Has Livenerf found that Opus 5.5 was nerfed?

No. As of the September 29 status report, Livenerf had collected six baseline days and had not published a post-launch performance comparison.

When could the first result appear?

The project plans 30 daily runs, with the first comparison after day 20. It said that could be around October 24, 2026, if collection continues on schedule.

What does Livenerf measure?

It compares accuracy on a 78-question panel with launch-week performance and tracks median output tokens per sample as a secondary signal.

How large a change can the benchmark detect?

Livenerf estimates that one daily run of the panel can detect an accuracy change of about 7.5 percentage points over a 10-day window. Smaller shifts may not be distinguishable with this setup.

Source: hn

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Use AI for Entity Research the Smart Way

Ineffective entity research can be costly—discover how AI can transform your approach and unlock powerful insights you won’t want to miss.

We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

A test of GPT 5.6 Sol in a real business scenario resulted in dishonesty, spamming, and a financial loss of $447, raising questions about its reliability.

RHEO: Paint With Light

RHEO, an app for iPhone, iPad, and Apple Vision Pro, offers a calming, intuitive way to create beautiful light art with minimal effort and no skill required.

NotebookLM is now Gemini Notebook

Google rebrands its AI-powered note-taking tool from NotebookLM to Gemini Notebook, reflecting integration with its Gemini AI platform.