Quantifying infrastructure noise in agentic coding evals

Anthropic engineers report that infrastructure configuration swings agentic coding benchmark scores by up to six percentage points. Internal experiments on Terminal-Bench 2.0 reveal that resource enforcement methods significantly alter results. This finding challenges the precision of current leaderboards and suggests developers should treat these scores with caution when comparing model capabilities.

Cover image for Quantifying infrastructure noise in agentic coding evals