Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts
Names benchmark reuse during automatic harness optimization as the evaluation risk.
3 items · 3 sources · 2 days
Operational story trace
Follow in this browser to see new updates on your Live feed.
Latest change
Two Sep 30 arXiv papers extend automatic harness evolution to robot-agent interfaces and to lifelong learning from research, after a Sep 16 paper flagged benchmark overfitting as its core evaluation risk.
On Sep 16 an arXiv paper warned that automatic harness optimization, which tunes prompts, memory, retrieval, tools, and control code against a released benchmark, can overfit that benchmark. Two more papers on Sep 30 treat harness evolution as a lever to be scaled and made lifelong.
Arc
Names benchmark reuse during automatic harness optimization as the evaluation risk.
Applies automatic harness evolution to a coding agent used as a robot policy via screenshots and a virtual gripper.
Proposes evolving the agent harness continually, drawing on published research.
Names benchmark reuse during automatic harness optimization as the evaluation risk.
Applies automatic harness evolution to a coding agent used as a robot policy via screenshots and a virtual gripper.
Proposes evolving the agent harness continually, drawing on published research.
What to watch — open questions
Storylines are threaded mechanically from the feed: stories that share a distinctive anchor across multiple days and sources. Each item links to its original source. The evidence trace, current state, and open questions are written by the editor routine and refreshed whenever a new beat lands.