Introducing advanced tool use on the Claude Developer Platform
Claude can now discover, learn, and select tools at runtime instead of working from a fixed toolset defined upfront.
17 articles · 4 categories
The finishable daily brief
Friday, Aug 28, 2026
17 articles · 4 categories
read top to bottom · then stop
In 30 seconds
Anthropic pushed four engineering posts today covering how Claude discovers and uses tools, how Claude Code runs sandboxed, and how to measure agent evals honestly — the clearest single-day view yet into how the company builds and tests its own agents.
The other big story: Z.ai confirmed its unbranded "Ox Alpha" leaderboard model was GLM-5.3-Flash all along, open-weighted it under a license aimed at hyperscalers, and revealed it runs on Chinese chips — while Qwen landed on a near-identical architecture independently and Tencent claims to have already beaten it.
Anthropic gave Claude dynamic tool discovery and gave Claude Code a sandboxed execution mode, the same day independent builders shipped a security proxy and a local memory layer for coding agents.
Claude can now discover, learn, and select tools at runtime instead of working from a fixed toolset defined upfront.
Filesystem and network isolation for Claude Code cuts permission prompts while limiting what a misbehaving run can touch.
A new hardware standard from Anthropic gives agents a common interface for taking action outside the browser and terminal.
A new open proxy sits between coding agents and the systems they touch, enforcing policy on the commands and network calls that get through.
An open local-first memory layer for coding agents reports 96% recall on the LongMemEval benchmark with no hosted backend.
Anthropic published two posts on measuring agent evals honestly, while outside builders reported hitting — and in one case cutting through — the same reliability ceiling in production coding agents.
Anthropic measures how much eval-score variance comes from infrastructure flakiness rather than real model differences, and what to control for.
A practical breakdown of what makes agent evals different from single-turn LLM evals, and where teams typically get them wrong.
An engineer argues current coding agents hit a reliability ceiling well short of full autonomy, and lays out where the gap actually sits.
Swapping a static tool list for dynamically loaded tools cut context overhead by more than 80% in one builder's coding agent.
Z.ai confirmed its unbranded "Ox Alpha" leaderboard model was GLM-5.3-Flash, open-weighted it under a license aimed at hyperscalers, and revealed it runs on Chinese chips — the clearest sign yet that Chinese labs are converging fast on cheap, fast flash-model architectures.
Z.ai open-weighted GLM-5.3 under new license terms written specifically to restrict large cloud providers from repackaging it.
The unbranded "Ox Alpha" model that topped leaderboards was GLM-5.3-Flash all along, trained and served on domestic Chinese silicon.
Z.ai and Alibaba's Qwen team landed on near-identical flash-model architectures independently, pointing to a shared design consensus for cheap, fast serving.
Tencent claims its newest model beats Z.ai and Moonshot on benchmarks, adding a third contender to the day's Chinese-lab model race.
Meta extended its custom-silicon strategy into networking hardware, and AWS/Databricks shipped infrastructure updates for feature stores, time-series forecasting, and PyTorch training at production scale.
Meta detailed MTIA 300, its next in-house accelerator for training ranking and recommendation models, part of a custom-silicon push now extending into networking hardware.
Decathlon deployed AWS's Chronos-2 time-series model to forecast weekly demand for tens of thousands of products across multiple continents.
SageMaker Feature Store added BatchWriteRecord, writing up to 25 records across feature groups in one call, plus ListRecords for enumerating record IDs.
Databricks details how it optimizes for "goodput" — useful training throughput net of failures — in its managed PyTorch training runtime.
You are caught up for this edition