{"schema_version":1,"kind":"agent-skill-lab","id":"lab-protocol","slug":"protocol","pilot_edition":0,"pilot_size":3,"state":"protocol","date":"2026-09-04","generated_at":"2026-09-04T00:00:00Z","featured_until":"2026-09-18","title":"Agent skills need receipts","question":"What changes when the same coding agent gets no skill, a short checklist, or the complete skill?","summary":"Agent Skill Lab will compare repeated runs of one fixed repository task, then publish the trajectory evidence behind the verdict.","method":{"task":{"description":"Each result uses one bounded repository task with a real failure or behavior change and a clean starting revision.","success_criteria":["The requested behavior passes deterministic end-to-end and focused tests.","The patch avoids unrelated product or repository changes.","The final explanation matches the observable implementation and test evidence."],"evaluation_method":"Score every run against the same predeclared success and quality rubric, then compare outcomes and observable execution traces."},"environment":{"model":{"name":"Pinned before each result","version":"Disclosed with the result"},"reasoning_effort":"Pinned before each result","harness":{"name":"Pinned before each result","version":"Disclosed with the result"},"repository_fixture":{"name":"Public or publication-safe fixture","revision":"Pinned before each result"},"permissions":["The same sandbox and tool permissions in every condition"],"budget":{"timeout_seconds":1800,"max_tokens_per_run":100000,"max_cost_usd_per_run":5}},"runs_per_condition":3,"conditions":[{"id":"no-skill","label":"No skill","setup":"The agent receives the task and normal harness instructions only."},{"id":"minimal-instructions","label":"Minimal instructions","setup":"The agent also receives a short checklist that captures the skill's main behavioral claim."},{"id":"full-skill","label":"Full skill","setup":"The agent receives the complete skill at one pinned revision."}],"held_constant":["Model, version, reasoning effort, and harness","Repository fixture, starting revision, task, and success criteria","Tool permissions, timeout, token ceiling, and cost ceiling","Number of independent runs in each condition"],"measures":[{"id":"task-success","label":"Task success","description":"Whether the run clears every predeclared success criterion."},{"id":"final-quality","label":"Final quality","description":"A zero-to-100 score from the same disclosed rubric for every run."},{"id":"trajectory","label":"Trajectory","description":"Planning, tool use, recovery events, interventions, and unnecessary actions."},{"id":"efficiency","label":"Efficiency","description":"Elapsed time, input and output tokens, tool calls, and estimated cost."}]},"publication_rule":"No verdict appears until all nine runs, the scoring receipts, and the publication-safe artifacts pass review.","limitations":["Three runs per condition expose obvious variance but cannot establish a universal ranking.","Each result applies to its pinned task, model, harness, and skill revision.","Published traces show observable actions and outputs, not hidden model reasoning."]}