Supabase Evals: Benchmark for testing how well AI agents build using Supabase
Supabase's first evals benchmark scores how well AI agents build correctly against a live Supabase backend, giving builders a concrete signal beyond vibes.
10 articles · 3 categories
The finishable daily brief
Saturday, Aug 1, 2026
10 articles · 3 categories
read top to bottom · then stop
In 30 seconds
China's frontier labs pushed harder on price and scale: DeepSeek cut its agentic output floor to $0.28 and is building an Inner Mongolia data center, while Moonshot's Kimi K3 lifted its valuation to $35B. Washington answered with two new export-enforcement threads — a claim that Alibaba's Nvidia chip use violated procurement rules, and a push to frame AI model distillation as a security issue.
Agent builders also got new tooling: three coding agents shipped (DSCode, Jcode, a sandboxed macOS agent), Supabase launched its first agent evals benchmark, and OpenAI published results on ten open math/CS problems.
Three independent coding agents shipped alongside the first dedicated evals benchmark for agents that build against a real backend — the tooling and measurement layer for agent-driven development kept expanding in parallel.
Supabase's first evals benchmark scores how well AI agents build correctly against a live Supabase backend, giving builders a concrete signal beyond vibes.
DeepSeek-powered terminal coding agent that plans, edits, tests, and reviews code, with local sandboxed sessions and visible token costs.
New terminal coding agent pitched as outperforming OpenCode and Pi — one more entrant in an increasingly crowded field.
Runs the Pi coding agent inside an Apple container on macOS with zero npm installed on the host — a sandboxing pattern for keeping agent tooling off the base system.
OpenAI published genuinely new results on long-standing open math and theoretical-CS problems — a research capability signal beyond typical benchmark leaderboards.
OpenAI published solutions to ten previously open problems spanning geometry, cryptography, and computational complexity, evidence frontier reasoning models are contributing genuinely new math.
Chinese labs kept undercutting on price and scaling compute — DeepSeek's $0.28 output floor and new Inner Mongolia data center, Moonshot's $35B valuation on Kimi K3 — while Washington opened fresh export-enforcement fronts over Alibaba's Nvidia chip use and AI distillation.
Moonshot AI's valuation reached $35B on the strength of Kimi K3, the latest sign Chinese labs are drawing serious investor capital.
DeepSeek dropped its agentic API output price to a $0.28 floor, the most aggressive pricing yet among frontier-capable Chinese labs.
Training smaller models on a rival's outputs, AI distillation is now being framed as a US-China policy flashpoint rather than just a technical debate.
DeepSeek is building a large new AI data center in Inner Mongolia, backing its aggressive pricing push with dedicated compute capacity.
Washington claims Alibaba's use of Nvidia chips to power Moonshot's Kimi K3 violated export-procurement rules, opening a new enforcement front against Chinese AI compute supply chains.
You are caught up for this edition