⬢github Python · 258 ★ +87 since we first saw it · pushed 6 h ago
ninjahawk/livenerf
Benchmark for tracking model capability after release.
livenerf is a 30-day, append-only benchmark testing whether Claude Opus 5.5 quietly degrades ('gets nerfed') after its 2026-09-22 launch. It runs daily via headless Claude Code on a frozen panel of 78 hard questions selected for inconsistent results, tracks drift statistically using Inspect AI and Anthropic's error-bars methodology, and logs everything raw. First results are expected around day 20.
Why now: It's getting attention because it addresses widely circulated community claims of post-launch model degradation with a pre-registered, day-0-baseline experiment instead of anecdote — and the daily series is actively running right now.
Who it is for: AI researchers, eval engineers, and anyone tracking whether frontier model quality changes after release.
Stars over our 10 snapshots: 171 to 258, since 2 h ago.
Where people talked about it
- Yhn Livenerf: Has Opus 5.5 been nerfed yet? 11 min ago
API: https://socialmediatrends-api.osmike.com/v1/repos/ninjahawk/livenerf