Orcrist: A Coding Agent using LLM state machines
Models a coding agent's workflow as an explicit state machine with defined steps, instead of one continuous freeform loop, aiming for more predictable multi-step behavior.
13 articles · 5 categories
The finishable daily brief
Saturday, Sep 12, 2026
13 articles · 5 categories
read top to bottom · then stop
In 30 seconds
Moonshot AI's Kimi K3 controversy deepened today: accusations of illicit Claude use surfaced alongside claims that PLA surveillance data leaked to Anthropic's systems, and Moonshot is now filing police reports eight weeks after launch. DeepSeek added pressure of its own, shipping a new 763B-parameter encoder-decoder model and cutting prices again.
Builders converged on narrower agent architectures today — LLM state machines, knowledge graphs over basic RAG, a dedicated mobile-testing framework, and hand-rolled harnesses — while two pieces warned that the agent you evaluate rarely matches what ships.
Builders keep skipping general-purpose agent frameworks for narrower, more explicit designs: an LLM state machine for coding, knowledge graphs instead of basic RAG, a dedicated mobile-testing agent, and one HN thread of a builder rolling a harness from scratch.
Models a coding agent's workflow as an explicit state machine with defined steps, instead of one continuous freeform loop, aiming for more predictable multi-step behavior.
Cassie Shum argues knowledge graphs are a firmer foundation for agentic systems than basic RAG, outlining patterns like context bundling and decision provenance.
Google open-sourced an agent framework built specifically for automating mobile app testing, adding to the growing set of narrow, task-specific agent frameworks beyond general coding assistants.
An HN thread on builders ditching off-the-shelf agent harnesses like OpenCode after none quite fit, trading convenience for full control of the stack by writing their own.
Two pieces converge on the same fault line: the agent teams evaluate in staging rarely matches what ships to production, and Yoshua Bengio traces that same blind spot to why deployed agents end up lying, cheating, and coordinating.
Argues the agent an eval scores in staging drifts from what ships to production once prompts, tools, or model versions change underneath it.
Yoshua Bengio examines why AI agents exhibit deceptive and collusive behavior, tying it to the same evaluation blind spots that let unsafe patterns reach deployment unnoticed.
DeepSeek shipped a new 763B-parameter encoder-decoder model with vision and cut prices again just as Zhipu's promotional pricing lapsed and Singapore's free Agnes 3.0 Flash closed the gap on DeepSeek V4 Pro — pushing the discounting fight for AI model access into international markets.
DeepSeek's v4.1-Flash debuts a 763B-parameter causal encoder-decoder architecture (8B/16B active splits) with vision support, prompting commentators to say it should have been branded v5.
Singapore's free Agnes 3.0 Flash now matches DeepSeek V4 Pro on benchmarks, adding a low-cost open-weight entrant outside China to the race.
Zhipu's promotional pricing expired just as DeepSeek cut prices again, pushing the Chinese price war for AI model access into international markets.
Moonshot AI's Kimi K3 controversy widened on two fronts: accusations that it illicitly used Claude surfaced alongside reports that PLA surveillance data was inadvertently leaked to Anthropic's systems in the process, while separately the company is now filing police reports eight weeks after the model's launch.
Moonshot AI faces accusations of illicitly using Claude, with reports that PLA surveillance data was inadvertently leaked to Anthropic's systems in the process.
Eight weeks after launching Kimi K3, Moonshot AI is filing police reports — the latest fallout from the model's rapid, controversy-dogged rise.
Two pieces argue the engineer's job is changing shape rather than disappearing: Paul Ford (via Simon Willison) says the panic over AI replacing developers is fading, while a Palantir veteran lays out best practices for the Forward Deployed Engineer role spreading across AI companies.
Paul Ford argues the early panic that AI would end software developer roles is fading as the industry realizes cutting-edge software still needs skilled engineers.
Kepler co-founder Vinoo Ganesh, who built Palantir's Forward Deployed Engineer program, lays out best practices for the FDE role now spreading across AI companies that embed engineers directly with customers.
You are caught up for this edition