Kimi K3 now available for public download
Moonshot's 2.8-trillion-parameter model is downloadable directly, following its initial launch.
20 articles · 6 categories
The finishable daily brief
Monday, Jul 27, 2026
20 articles · 6 categories
read top to bottom · then stop
In 30 seconds
Moonshot's 2.8-trillion-parameter Kimi K3 became publicly downloadable today, and vLLM and Modal already have day-0 serving support live — vLLM ships hybrid KDA prefix caching and DSpark speculative decoding, Modal pairs the model with a custom-trained DFlash speculator. A HackerNoon portability audit pushed back on the launch-week hype, arguing K3's headline benchmark scores don't hold up once you account for how it was evaluated.
Two new tools target unsafe agent tool calls directly: Belay adds a local firewall for coding agents, and a static verifier built on the "Guardians of the Agents" formal-verification paper checks OpenCode's tool calls before they run. Separately, a widely syndicated report put Chinese models at nearly a third of enterprise LLM tokens at one-tenth the cost of US rivals, even as DeepSeek paused a new funding round.
Moonshot's 2.8-trillion-parameter Kimi K3 is now publicly downloadable with day-0 vLLM and Modal serving support, but a portability audit is already challenging its launch-week benchmark claims.
Moonshot's 2.8-trillion-parameter model is downloadable directly, following its initial launch.
HackerNoon argues K3's headline benchmark numbers don't hold up once you account for how the model was evaluated.
vLLM's Kimi K3 support includes hybrid KDA prefix caching and DSpark speculative decoding, tuned for both NVIDIA and AMD GPUs.
Modal's Kimi K3 deployment pairs the 2.8T-parameter model with a custom-trained DFlash speculative-decoding model.
A case study shows a coding agent refactoring a 750,000-line codebase over three days with no human code review, while GitHub frames its own harness, not any single model, as the unit that makes agentic workflows reliable.
The agent rebuilt a core system invariant in three days, running 31 verification passes and correcting 201 errors before shipping with no human review.
GitHub's Copilot workflow guidance argues a consistent prototype-plan-implement-review harness matters more than chasing each new model release.
Hyre chains an LLM pipeline (PDF-to-markdown, layout analysis, data extraction, scoring) to turn a resume into a numeric evaluation.
Two new builder tools, a local firewall and a static verifier, check agent tool calls before execution, while a new benchmark shows single-turn evals miss the multi-turn regressions that actually break agents.
Belay intercepts and filters coding-agent actions locally, before they reach the file system or shell.
Built on the "Guardians of the Agents" formal-verification paper, the plugin statically checks OpenCode's tool calls before execution.
Single-turn benchmark scores overstate reliability; the real failure mode is regressions across a persistent workspace, not missing features.
The alliance frames open-source security practices as infrastructure for AI safety, extending the model that secured cloud and telecom software.
Netflix detailed its in-house Triton/vLLM serving stack, LangChain built a sub-second full-text search index over agent traces, and AWS pitched task-aware compression as a fix for RAG's document-scale ceiling.
Netflix built its own serving layer on Triton and vLLM rather than relying on a third-party inference API.
SmithDB indexes large, deeply nested agent-trace JSON in object storage with a median 400ms search latency.
Traditional RAG breaks down on analytical tasks spanning hundreds of documents; AWS's TAKC pre-compresses knowledge bases into task-specific representations instead.
InfoQ argues traditional API gateways assume deterministic services, an assumption agentic AI breaks, pushing engineering leaders toward AI gateways as a dedicated architecture seam.
A widely syndicated report put Chinese models at nearly a third of enterprise LLM tokens at one-tenth the cost of US rivals, even as DeepSeek paused a new funding round and US lawmakers weigh restrictions that could slow the catch-up.
A widely syndicated report puts Chinese open models at roughly 30% of enterprise LLM token volume, priced at about one-tenth of comparable US models.
The pause follows leaked comments attributed to founder Liang Wenfeng that went viral days earlier.
Washington Examiner frames pending congressional measures as a risk to US competitiveness against China's AI pace.
Anthropic named Cognizant a Global Premier Partner after training over 30,000 of its associates on Claude, and new OpenAI research tracks how ChatGPT is reshaping what workers actually do inside their roles.
Cognizant has trained 30,000+ associates on Claude and is embedding it across its platforms under the expanded partnership.
The study finds ChatGPT users are taking on tasks that cross traditional role boundaries, not just automating existing ones.
You are caught up for this edition