Can Jev Be a Better Agent Evaluator?
LangChain benchmarked "Jev-as-a-Judge," a non-LLM classifier, against LLM judges for agent evals on accuracy, repeatability, latency, and cost.
14 articles · 4 categories
The finishable daily brief
Sunday, Sep 20, 2026
14 articles · 4 categories
read top to bottom · then stop
In 30 seconds
Four Chinese labs — Alibaba, DeepSeek, Kimi, and Step — shipped frontier model previews on the same day, pushing 1M-token context, KV-cache compression, and a 7B open-weight image model that claims to beat closed rivals.
On the tooling side, Unity, Alibaba, and independent developers all shipped competing coding-agent bets, while a Z.ai data-handling scandal and a Chinese model found running US Federal Register searches renewed questions about trusting vendor-controlled AI infrastructure.
Today's engineering-practice news was about how agents get judged and coordinated, not just what they can do.
LangChain benchmarked "Jev-as-a-Judge," a non-LLM classifier, against LLM judges for agent evals on accuracy, repeatability, latency, and cost.
Google's ADK for Kotlin hit 1.0 with full parity to the Python SDK, adding on-device AI support across Android, Kotlin, and JVM/server targets.
An open-source durable kernel (Python, MCP) aims to give multi-agent swarms persistent state and coordination instead of ad-hoc orchestration scripts.
Vendors and indie builders shipped competing bets on coding-agent distribution and speed — official IDE integrations, an open-sourced review agent, and two from-scratch agents chasing bootstrap simplicity and raw throughput.
Unity shipped first-party Claude Code and Codex plugins, folding coding-agent workflows directly into its editor instead of leaving them to third-party integrations.
Alibaba open-sourced OpenCodeReview, a CLI that pairs deterministic file selection and rule matching with an LLM agent for dynamic review analysis.
Cotyper launched claiming 5x the speed of Codex on coding tasks, betting raw throughput — not just capability — is the next coding-agent differentiator.
Bailout is a minimal coding agent built to bootstrap a fresh VM — no GitHub auth, no config, no API key — before your real agent tooling is set up.
Four Chinese labs pushed frontier releases on the same day, spanning open-weight image generation, long-context compression, and price-down pressure on frontier inference.
Alibaba's Qwen-Image-2.1 is a 7B-parameter open-weight image model that its maker claims outperforms larger closed alternatives.
DeepSeek-V4.1-Flash pairs a causal encoder-decoder MoE architecture with 1M-token context and aggressive KV-cache compression to cut long-context serving cost.
Kimi's K2.8 preview extends context to 1M tokens while keeping the same model ID, letting existing integrations pick up the upgrade without a version bump.
Step 5's preview is pitched as the cheapest frontier-class model out of China yet, continuing the price-down pressure on frontier inference.
Two stories cast doubt on trusting vendor-controlled AI infrastructure: Z.ai's encrypted, unauthorized uploads left users unable to verify deletion, and a Chinese model turned up inside a US government search system.
Z.ai uploaded a user's workspace with encryption only it controls, so the only assurance the data was deleted is Z.ai's own word.
The unauthorized-upload incident is drawing wider scrutiny of Z.ai's data-handling practices from Chinese and international press.
The US Federal Register was found running document search on a Chinese-made AI model, raising provenance and vendor-trust questions for government AI procurement.
You are caught up for this edition