LLM Digest
Subscribe

AI Daily Recap

18 articles · 4 categories

View as JSON

The finishable daily brief

What happened in AI — Jul 25, 2026

Saturday, Jul 25, 2026
18 articles · 4 categories

read top to bottom · then stop

In 30 seconds

  • Kimi K3's open weights ship Sunday, pitched on self-hosting to cut China data-routing risk — not on beating US models, which still score 76% vs Kimi's 32% on a cyber benchmark.
  • Anthropic's Opus 5 hits Fable-level performance at half Fable's price and, per its system card, is the least prompt-injectable Claude yet.
  • ActionRail ships as an open runtime for grounding agent actions against policy checks before execution.
  • LangChain: the durable moat is who owns the agent system and context pipeline, not the underlying model.
  • AWS released AWS-bench, an open benchmark for evaluating AI agents against AWS infrastructure.

Kimi K3's open weights land Sunday, and the real draw is self-hosting, not the scoreboard: US frontier models still more than double Kimi K3's score on an independent cybersecurity benchmark (76% vs 32%). Anthropic moved the price/performance line the same week — Opus 5 pairs Fable-level output at half Fable's price and, per its system card, is Anthropic's hardest model yet to prompt-inject.

On the building side, a new open-source action-grounding runtime and AWS's own agent benchmark both target the same problem: judging what agents actually do, not just how well they reason.

Kimi K3's Open Weights Land Sunday — Self-Hosting Is the Draw, Not the Score 2 items

Moonshot AI opens Kimi K3's weights this weekend so enterprises can self-host instead of routing data through a China-based API, even as independent testing shows US frontier models still more than double Kimi K3's score on a cybersecurity benchmark.

Opus 5's Price Cut Comes With a Prompt-Injection Number 2 items

Anthropic's Opus 5 lands at half Fable's price for Fable-tier output, and its system card shows the largest prompt-injection resistance gain of any Claude model to date.

Quoting Boris Cherny

simon_willisonJul 25Details

Anthropic's Boris Cherny highlights that Opus 5 is the least prompt-injectable Claude yet across PI evals and red-teaming — a detail buried in the system card rather than the headline benchmarks.

Builders Ship Grounding Tools and Rethink Who Owns the Agent Stack 3 items

New tooling and commentary converge on the same theme: agent reliability and defensibility come from grounding actions and owning the surrounding system, not from the model alone.

Evals Push Toward Context Engineering, Not Bigger Models 2 items

Two releases reinforce that agent evaluation is moving from raw model reasoning toward the pipelines and benchmarks that surround it.

You are caught up for this edition