{"slug":"harness-evolution","label":"Harness Evolution","item_count":3,"day_count":2,"source_count":3,"first_seen":"2026-09-16T09:21:33+00:00","last_updated":"2026-09-30T17:00:38+00:00","generated_at":"2026-10-01T15:03:55.097786+00:00","sources":["arxiv_agent_systems_research","arxiv_cs_ai","arxiv_llm_reliability"],"days":[{"date":"2026-09-16","items":[{"title":"Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts","url":"http://arxiv.org/abs/2609.18366v1","source":"arxiv_cs_ai","type":"paper","summary_1line":"Reliable agent evaluation is complicated by automatic harness optimization, which repeatedly uses a released benchmark $B_{\\mathrm{rel}}$ to guide a Proposer that edits prompts, memory, retrieval, tools, and control c...","why_it_matters":"Matches feed focus: agent, harness, evaluation.","sid":"38522ce275c55bf2","published":"2026-09-16T09:21:33+00:00","editor_note":"Names benchmark reuse during automatic harness optimization as the evaluation risk."}]},{"date":"2026-09-30","items":[{"title":"Scale and Selection: What Makes Automatic Harness Evolution Work for Visual-Interface Robot Agents","url":"http://arxiv.org/abs/2609.39304v1","source":"arxiv_llm_reliability","type":"paper","summary_1line":"When an off-the-shelf coding agent is used directly as a robot policy, observing a browser-based 3D interface through screenshots and acting by posing a virtual target gripper through a few tools, the agent's harness,...","why_it_matters":"Matches feed focus: agent, harness, evaluation.","sid":"50fe4e3ccee5061b","published":"2026-09-30T08:46:38+00:00","editor_note":"Applies automatic harness evolution to a coding agent used as a robot policy via screenshots and a virtual gripper."},{"title":"Learning from Research: Toward Lifelong Agent Harness Evolution","url":"http://arxiv.org/abs/2609.40169v1","source":"arxiv_agent_systems_research","type":"paper","summary_1line":"Language agents are expected to solve increasingly complex tasks, creating a growing need for continual improvement. One promising approach is to evolve the agent harness, the software that governs tool use, memory ma...","why_it_matters":"Matches feed focus: agent, harness, eval.","sid":"f8f8651e579c3c68","published":"2026-09-30T17:00:38+00:00","editor_note":"Proposes evolving the agent harness continually, drawing on published research."}]}],"editorial":{"tldr":"On Sep 16 an arXiv paper warned that automatic harness optimization, which tunes prompts, memory, retrieval, tools, and control code against a released benchmark, can overfit that benchmark. Two more papers on Sep 30 treat harness evolution as a lever to be scaled and made lifelong.","stale":false,"whats_new":"Two Sep 30 arXiv papers extend automatic harness evolution to robot-agent interfaces and to lifelong learning from research, after a Sep 16 paper flagged benchmark overfitting as its core evaluation risk.","why_it_matters":"If your agent harness (tools, memory, retrieval, control flow) is tuned automatically against a public benchmark, its score measures the tuning loop as much as the agent; hold out private evals before trusting gains.","take_for_builders":"If you tune prompts, tools, or memory automatically, keep a private eval set the optimizer never sees and re-score after each harness change; treat these papers as method references, not drop-in tooling.","status":{"state":"Research · early","tone":"rising","changed":"2026-09-30","detail":"Three arXiv papers in two weeks, no shipped product or independent replication yet."},"beats":[{"kicker":"CAVEAT","tone":"turn","headline":"Bad Genius paper: harness optimization on a released benchmark risks shortcut learning","summary":"Automatic harness optimization repeatedly reuses a released benchmark to guide a Proposer that edits prompts, memory, retrieval, tools, and control code.","sids":["38522ce275c55bf2"]},{"kicker":"NOW","tone":"now","headline":"Harness evolution scales to a robot-agent interface and toward lifelong improvement","summary":"One paper studies what makes automatic harness evolution work for a coding agent driving a browser-based 3D robot interface; the other evolves the harness continually from research.","sids":["50fe4e3ccee5061b","f8f8651e579c3c68"]}],"open_questions":["Do the Sep 30 harness-evolution methods hold up on a held-out benchmark, given the Sep 16 overfitting warning?","Does automatic harness evolution transfer from the robot-interface setting to ordinary coding or tool-use agents?","Does any open-source agent framework adopt automatic harness evolution as a built-in step?"],"generated_at":"2026-10-01T15:03:54+00:00"}}