LLM Digest
Subscribe

AI Daily Recap

20 articles · 6 categories

View as JSON

The finishable daily brief

What happened in AI — Jul 27, 2026

Monday, Jul 27, 2026
20 articles · 6 categories

read top to bottom · then stop

In 30 seconds

  • Moonshot's Kimi K3 (2.8T params) went public with day-0 vLLM and Modal serving support, though a HackerNoon portability audit disputes its launch-week benchmark scores.
  • Belay and a new OpenCode static verifier both launched to catch unsafe agent tool calls before they execute.
  • philschmid's EvoCode-Bench tests coding agents across 227 sequential rounds, showing single-turn scores overstate reliability.
  • A coding agent refactored a 750,000-line codebase over three days with zero human code review, running 31 verification passes to catch 201 errors.
  • Chinese models now handle nearly a third of enterprise LLM tokens at one-tenth the cost of US rivals, per a widely syndicated report, even as DeepSeek paused a new funding round.
  • Anthropic named Cognizant a Global Premier Partner after training over 30,000 of its associates on Claude.

Moonshot's 2.8-trillion-parameter Kimi K3 became publicly downloadable today, and vLLM and Modal already have day-0 serving support live — vLLM ships hybrid KDA prefix caching and DSpark speculative decoding, Modal pairs the model with a custom-trained DFlash speculator. A HackerNoon portability audit pushed back on the launch-week hype, arguing K3's headline benchmark scores don't hold up once you account for how it was evaluated.

Two new tools target unsafe agent tool calls directly: Belay adds a local firewall for coding agents, and a static verifier built on the "Guardians of the Agents" formal-verification paper checks OpenCode's tool calls before they run. Separately, a widely syndicated report put Chinese models at nearly a third of enterprise LLM tokens at one-tenth the cost of US rivals, even as DeepSeek paused a new funding round.

Kimi K3 Goes Public, and the Skepticism Starts Immediately 4 items

Moonshot's 2.8-trillion-parameter Kimi K3 is now publicly downloadable with day-0 vLLM and Modal serving support, but a portability audit is already challenging its launch-week benchmark claims.

Coding Agents Take On Bigger, Less-Supervised Work 3 items

A case study shows a coding agent refactoring a 750,000-line codebase over three days with no human code review, while GitHub frames its own harness, not any single model, as the unit that makes agentic workflows reliable.

New Tools Aim to Catch Unsafe Agent Actions Before They Run 4 items

Two new builder tools, a local firewall and a static verifier, check agent tool calls before execution, while a new benchmark shows single-turn evals miss the multi-turn regressions that actually break agents.

Platform Teams Build Serving and Data Layers Purpose-Built for Agents 4 items

Netflix detailed its in-house Triton/vLLM serving stack, LangChain built a sub-second full-text search index over agent traces, and AWS pitched task-aware compression as a fix for RAG's document-scale ceiling.

China's Model Cost Advantage Keeps Widening, Even as DeepSeek Wobbles 3 items

A widely syndicated report put Chinese models at nearly a third of enterprise LLM tokens at one-tenth the cost of US rivals, even as DeepSeek paused a new funding round and US lawmakers weigh restrictions that could slow the catch-up.

Enterprise Adoption Signals 2 items

Anthropic named Cognizant a Global Premier Partner after training over 30,000 of its associates on Claude, and new OpenAI research tracks how ChatGPT is reshaping what workers actually do inside their roles.

You are caught up for this edition