Evaluating Agents Beyond the First Prompt
EvoCode-Bench tests coding agents across 227 sequential rounds in a persistent workspace. Single-turn scores overstate reliability — regressions, not missing features, are the real bottleneck. Context & related coverage →