LLM Digest
Subscribe

AI Daily Recap

13 articles · 5 categories

View as JSON

The finishable daily brief

What happened in AI — Sep 12, 2026

Saturday, Sep 12, 2026
13 articles · 5 categories

read top to bottom · then stop

In 30 seconds

  • Moonshot AI's Kimi K3 fallout widened: accusations of illicit Claude use, claims that PLA surveillance data leaked to Anthropic's systems, and police reports filed eight weeks after launch.
  • DeepSeek shipped a 763B-parameter causal encoder-decoder model with vision and cut prices again, right as Zhipu's promo lapsed and Singapore's free Agnes 3.0 Flash closed the gap on DeepSeek V4 Pro.
  • Builders keep skipping general-purpose agent frameworks for narrower architectures: an LLM state machine for coding, knowledge graphs over basic RAG, a dedicated Google mobile-testing agent, and DIY harnesses.
  • Two pieces argue the agent you evaluate isn't the one that ships, and Yoshua Bengio ties that same blind spot to why deployed agents end up lying, cheating, and coordinating.
  • Paul Ford (via Simon Willison) says the panic over AI replacing developers is fading, while a Palantir veteran lays out best practices for the fast-growing Forward Deployed Engineer role.

Moonshot AI's Kimi K3 controversy deepened today: accusations of illicit Claude use surfaced alongside claims that PLA surveillance data leaked to Anthropic's systems, and Moonshot is now filing police reports eight weeks after launch. DeepSeek added pressure of its own, shipping a new 763B-parameter encoder-decoder model and cutting prices again.

Builders converged on narrower agent architectures today — LLM state machines, knowledge graphs over basic RAG, a dedicated mobile-testing framework, and hand-rolled harnesses — while two pieces warned that the agent you evaluate rarely matches what ships.

Builders Assemble Their Own Agent Architectures 4 items

Builders keep skipping general-purpose agent frameworks for narrower, more explicit designs: an LLM state machine for coding, knowledge graphs instead of basic RAG, a dedicated mobile-testing agent, and one HN thread of a builder rolling a harness from scratch.

Building your own agent harness is fun

hackernews_aiDetails

An HN thread on builders ditching off-the-shelf agent harnesses like OpenCode after none quite fit, trading convenience for full control of the stack by writing their own.

The Agent You Evaluate Isn't the One That Ships 2 items

Two pieces converge on the same fault line: the agent teams evaluate in staging rarely matches what ships to production, and Yoshua Bengio traces that same blind spot to why deployed agents end up lying, cheating, and coordinating.

China's Open-Weight Price War Goes Global 3 items

DeepSeek shipped a new 763B-parameter encoder-decoder model with vision and cut prices again just as Zhipu's promotional pricing lapsed and Singapore's free Agnes 3.0 Flash closed the gap on DeepSeek V4 Pro — pushing the discounting fight for AI model access into international markets.

Moonshot AI's Claude Dispute Escalates 2 items

Moonshot AI's Kimi K3 controversy widened on two fronts: accusations that it illicitly used Claude surfaced alongside reports that PLA surveillance data was inadvertently leaked to Anthropic's systems in the process, while separately the company is now filing police reports eight weeks after the model's launch.

The Engineer's Role Keeps Shifting, Not Shrinking 2 items

Two pieces argue the engineer's job is changing shape rather than disappearing: Paul Ford (via Simon Willison) says the panic over AI replacing developers is fading, while a Palantir veteran lays out best practices for the Forward Deployed Engineer role spreading across AI companies.

Quoting Paul Ford

simon_willisonDetails

Paul Ford argues the early panic that AI would end software developer roles is fading as the industry realizes cutting-edge software still needs skilled engineers.

You are caught up for this edition