A 640-point Hacker News launch has outrun Livenerf's evidence. The public benchmark was built to test whether Claude Opus 5.5 gets worse after release, yet its latest progress update covered only six of 30 planned daily runs. All six belonged to the baseline. The project's first possible verdict is around October 24, and calling a nerf before then would break the rules that make the experiment worth watching. The repository's status page says exactly where the clock stands.
That gap between attention and evidence is the story. Model users regularly feel that a familiar system has become shorter, less capable or less willing to finish a task. A screenshot or one failed coding session cannot separate a model change from randomness, a changed prompt, a client update, safety routing or a bad afternoon. Livenerf lists those competing explanations as the reason for freezing everything it can measure.
The useful part is the delay
Livenerf's decision rule was written before the time series began. Days 1 through 10 establish the baseline. Days 11 through 20 form the first comparison window, and days 21 through 30 form the second. A change counts only if the 99% confidence interval excludes zero in both later windows, the movement goes in the same direction, and each shift is at least three percentage points. The harness must also retain the same Claude Code version and sample-shaping hash, while errors must stay below 5%. The pre-registration says every lesser result will be reported as "no change detected".
This makes day 30 more than a calendar target. It prevents an eye-catching daily dip from becoming a finding. The repository says the first published results row will arrive after day 20, but even that row cannot satisfy the two-window rule on its own. A test designed around delayed judgment is a useful counterweight to a discussion that reached 640 Hacker News points while the baseline was still being assembled. Livenerf's generated design puts the earliest possible call at day 30.
The panel is built from Opus 5.5's mistakes
The author screened 2,336 questions drawn from GPQA Diamond, a fixed 2,000-question MMLU-Pro subset, competition mathematics and AIME 2025-26. Each question received four samples. Opus 5.5 answered about 93% correctly on the first try, and 2,119 questions were always right while 133 were always wrong. Those questions reveal little about a small change in ability, so the final panel keeps 78 questions that the model answers correctly only some of the time. The calibration report publishes the yield for each benchmark.
There is a subtle correction behind that selection. Choosing questions because they land near the middle makes their screening scores look closer to 50% than they may be. Livenerf therefore ran eight fresh confirmation samples for every eligible question and used those new rates only for its power calculation. Across the confirmed set, the mean pass rate rose from 54.7% during screening to 62.0% during confirmation. The calibration record labels this as selection bias rather than model improvement.
At one pass per question per day, the projected minimum detectable effect is 7.5 accuracy points for a 10-day window, at 80% power under the pre-registered 99% test. Each daily run contains 78 Opus 5.5 samples plus 12 control samples, for 90 in total. The project's September 29 update reported six complete days, no misses, one harness hash and Claude Code 2.1.280 throughout. The design estimates that schedule at 3.6 points of the author's weekly plan meter.
The detector has already tested itself
Before starting the baseline, Livenerf ran 1,248 graded validation samples across four interleaved arms: Opus 5.5 at high, medium and low effort, plus Opus 5 at high effort. The low-effort arm used 62% fewer output tokens than high effort and lost 8.3 accuracy points, with a standard error of 4.5 points. Medium effort used 26% fewer tokens and lost 4.2 points, with a 3.9-point standard error. The validation passed its pre-registered check because the token drop for low effort excluded zero at 99%.
The same exercise exposed a limit. Replacing Opus 5.5 with Opus 5 produced a measured accuracy difference of minus 3.8 points with a 6.3-point standard error. Output tokens fell 23%, but the 99% interval ran from minus 46% to plus 8%. The instrument could not distinguish that same-family swap at its validation sample size. Livenerf discloses the miss prominently, which matters because a smaller model behind the same name is one of the explanations users offer for a perceived nerf. The project's own validation calls that swap indistinguishable.
Its A/A check deserves the same restraint. Splitting repeated high-effort samples produced a 6.4-point difference with a 3.6-point standard error, giving a z-score of 1.79. That remained inside the pre-registered acceptance boundary of 1.96, though it sat close enough to show how much ordinary variation the series must absorb. A passed validation means the detector caught its chosen positive control. It does not mean every serving change will be visible. The full validation table reports the interval alongside the effort tests.
What Livenerf is actually observing
The target is Opus 5.5 as delivered through headless Claude Code on a Max subscription. It is not a direct measurement of the raw Anthropic API model. Every sample uses explicit high effort, a frozen system prompt and an empty working directory. Tools, MCP servers, settings, hooks, project instructions and memory are removed. If Claude Code retries a turn, serves another model or ends in a refusal, the sample is rejected and logged as a classifier event instead of being scored. Those boundaries are part of the pre-registered definition of the experiment.
That scope matches a real user path, but it bundles the model with Anthropic's serving and Claude Code layers. Anthropic says Opus 5.5 launched on September 22 with production safeguards enabled, and some cybersecurity, biology and frontier-model tasks can be handled by fallback models. The company also presents the release as faster and less expensive than Opus 5. Anthropic's launch post says default workloads cost 40% less in its tests. Livenerf's rejection rules are meant to stop visible fallbacks from contaminating its score, though an unobserved platform change could still affect the path.
The control arm helps with that attribution problem. Each day, Opus 5 runs the same 12 GPQA items under the same harness. If Opus 5.5 and the control shift in the same direction strongly enough, the pre-registration labels the result a harness or platform change rather than an Opus 5.5 change. The pairing follows a broader statistical argument: eval reports should quantify noise and compare like questions with like questions instead of treating one aggregate score as exact. Evan Miller's paper describes methods for measuring model differences and planning eval experiments.
The implementation sits on Inspect, the evaluation framework developed by the UK AI Security Institute and Meridian Labs. Inspect supplies datasets, solvers, scorers, logs and sandbox support. Livenerf adds its Claude Code provider, frozen panel, scheduling and decision logic. Inspect's documentation also supports external agents such as Claude Code and Codex CLI. That foundation makes the repository easier to examine than a dashboard backed by unpublished prompts.
The questions leave a large caveat
Seventy-eight questions are a thin slice of model behavior, and this slice favors exact answers in science, knowledge and competition math. It does not directly test a week-long refactor, a changing repository, writing quality or tool use. Livenerf can detect movement on its panel without proving that developers will notice the same movement in their work. The pre-registration defines the primary arm as those four benchmark families.
The items themselves are imperfect. A report-only audit covered the 78 panel questions and two later exclusions, classifying 30 as ambiguous and eight as having suspect answer keys. The audit did not remove them after the fact. A planned sensitivity analysis will recalculate the result without those items. There is another complication: Claude performed the audit, creating a stated conflict because Claude belongs to the model family under examination. The pre-registration records both the suspect items and the conflict.
Developers considering reuse should also notice the repository's license field. The code is public, but the README says a license has not yet been chosen. That means teams should inspect and discuss reuse terms before incorporating the package into their own systems. The current repository states "Not yet chosen" under License.
The result to watch
Livenerf will become informative through continuity. The first check is whether all 30 daily runs retain the same CLI, prompt, panel and harness hash. After day 20, the first comparison can show a direction and an interval. Around October 24, the second window can either confirm that movement, contradict it or leave the test below its 7.5-point detection range. The pre-registration requires both windows and keeps token counts secondary.
The 640-point launch shows that many users want a meter for model drift. Livenerf's better contribution is a rule for when that meter must stay silent. Watch the two 10-day deltas, the control arm, the error rate and any recorded deviation from the frozen harness. If October 24 arrives without a qualifying change, the honest result will be limited but useful: this test did not detect one. That is the null-result language the project committed to before collection.