Is Claude Dumb Today?

Daily HumanEvalPlus-CC164 benchmark for Claude Code (Opus 5)

...

Loading latest results…

Score
 
Model
 
Cost
 
Runtime
 

Score History (last 90 runs)

Where the models disagree

Tasks where Opus 5 and Opus 4.8 have different pass rates over recent paired runs. Green = always passes, red = always fails. Spread is the gap between the best and worst model on that task — a high spread reveals a real tradeoff, not noise. Historical divergences include HumanEval/97 (Python signed-modulo quirk) and HumanEval/141 (Unicode .isalpha() vs literal a–z range).

Task Opus 5 Opus 4.8 Spread
Loading…

Per-Task Results (latest run)

Task Function Result Base EvalPlus Attempts Turns Cost Error
Loading…