LLM Digest
Subscribe

AI Storyline

2 items · 2 sources · 2 days

View as JSON

Operational story trace

Claude's build-eval and hill-climb tooling

Latest change

Hamel Husain reviewed the new build_eval and hill-climb commands on Sep 30, calling a first-party Anthropic eval tool likely to shape how teams build evals.

Earlier contextThe story so far

Anthropic added build-eval and hillclimb commands to its claude-api plugin on Sep 28, alongside a post on designing evals without fooling yourself. Hamel Husain reviewed the tooling two days later.

editor-curated · source-linked

Arc

Sep 28Sep 30 · now
LAUNCH · Sep 28
Anthropic adds build-eval and hillclimb commands to the claude-api plugin
1 source · scout · show source ▾
REACTION · Sep 30
Hamel Husain reviews the tool: build evals, check graders, improve against them
1 source · scout · show source ▾

What to watch — open questions

  • Does the grader check in build_eval catch graders that score the wrong behavior, in independent use?
  • Do hillclimb gains hold up on held-out cases rather than the eval set it optimized against?
How this thread was built
scout surfaced 2editor wrote the arc · 2 beats

Storylines are threaded mechanically from the feed: stories that share a distinctive anchor across multiple days and sources. Each item links to its original source. The evidence trace, current state, and open questions are written by the editor routine and refreshed whenever a new beat lands.