{"date":"2026-07-22","title":"What happened in AI — Jul 22, 2026","generated_at":"2026-07-22T21:15:00Z","intro":["Moonshot's Kimi K3 forced a simultaneous economic and political reckoning today: vLLM shipped production-grade serving support, Microsoft is evaluating swapping ChatGPT and Claude out of Copilot to cut inference costs by $600M, and the White House escalated claims that the model was built on smuggled Nvidia chips and cloned Anthropic IP.","On the practice side, LangChain shipped a skill that generates evals straight from an agent's repo and production traces, and Anthropic published the containment architecture — filesystem, network, and execution limits — it uses to sandbox Claude across Web, Code, and Cowork."],"highlights":["vLLM shipped production-scale serving support for Kimi K3 while Microsoft weighs swapping it into Copilot to save $600M on inference.","The White House escalated accusations that Moonshot AI trained Kimi K3 on smuggled Nvidia chips and IP stolen from Anthropic.","LangChain's new Eval Engineering Skill turns an agent's repo and traces into runnable evals automatically.","Anthropic published the sandboxing architecture — filesystem, network, and execution limits — it uses to contain Claude across its own products.","A wave of Show HN coding-agent tools launched: a self-hosted LLM router (Millwright), two native-Mac agent terminals (Rabbitty, Forkbench), and an autonomous eval-writing 'AI engineer' (Langy).","Anthropic committed $200M to an Economic Futures Research Fund; OpenAI launched Presence, an enterprise voice/chat agent platform."],"article_count":19,"categories":[{"name":"Agent Engineering: Evals and Composition","slug":"agent-engineering-evals-composition","summary":"Eval generation and agent architecture both moved from talk to shipped tooling: LangChain now automates eval creation from real traces, and a conference talk argues agents need versioned, composable \"virtual tools\" instead of ad hoc prompt chains.","articles":[{"title":"Eval Engineering Skill: Build Evals From Repo Context and Traces","summary":"LangChain's new skill reads an agent's repo and production traces, proposes evals through user interviews, and outputs runnable Harbor test tasks.","source":"langchain_blog","url":"https://www.langchain.com/blog/towards-automating-eval-engineering","published":"Wed, 22 Jul 2026 17:18:03 GMT"},{"title":"3 Years of Graph Engineering with LangGraph","summary":"LangChain frames LangGraph's three years as validation that graph-based orchestration, not single-model prompting, is the durable pattern for reliable agents.","source":"langchain_blog","url":"https://www.langchain.com/blog/3-years-of-graph-engineering-with-langgraph","published":"Wed, 22 Jul 2026 12:37:19 GMT"},{"title":"Presentation: From Copy-Paste to Composition: Building Agents Like Real Software","summary":"Jake Mannix argues agents should move past ad hoc \"1970s BASIC\" architectures toward an intermediate protocol layer of versioned, encapsulated \"virtual tools.\"","source":"infoq_ai_ml","url":"https://www.infoq.com/presentations/agent-software-engineering/?utm_campaign=infoq_content&utm_source=infoq&utm_medium=feed&utm_term=AI%2C+ML+%26+Data+Engineering","published":"Wed, 22 Jul 2026 11:57:00 GMT"}]},{"name":"Security & Containment","slug":"security-containment","summary":"Agent containment moved from research topic to operational priority: Anthropic detailed the sandboxing limits it runs Claude under, a Chinese lab's model was reportedly hacked after copying OpenAI outputs, and analysts flagged cybersecurity as a rising theme across this week's AI coverage.","articles":[{"title":"Anthropic Details How It Contains Claude Across Web, Code, and Cowork","summary":"Anthropic argues agent safety depends on deterministic limits on an agent's filesystem, network, and execution environment, not just model-level alignment.","source":"infoq_ai_ml","url":"https://www.infoq.com/news/2026/07/anthropic-claude-containment/?utm_campaign=infoq_content&utm_source=infoq&utm_medium=feed&utm_term=AI%2C+ML+%26+Data+Engineering","published":"Wed, 22 Jul 2026 12:25:00 GMT"},{"title":"China's Zhipu AI model contains hack after OpenAI models go rogue - South China Morning Post","summary":"A report says Zhipu AI's model was compromised after ingesting outputs from misbehaving OpenAI models, an early example of contamination risk from cross-model data reuse.","source":"search_cn_open_weight_labs","url":"https://news.google.com/rss/articles/CBMizAFBVV95cUxQaGx5UU00ZVdpY0NlWi1HTTdhUVctNjdEREV4RzJyV2pLbDU2ZVVDdlJsRjlOMXEwLTE2SXhlbzdjcDBVNEd6cW05U25uN0xmYVlCT2RXZ3pCTUlvM0pjMEZ0VlRpQWRoeG8yWWxzUHBXQkF0TTNOcjlyaTlnRC1xMU5HTGU1anQ0czB6b2lMQnVGZTBYeHV4Y3FGcVBiTG93VFBGdjl4aTdlTGRmcngtRUMwZm1oMktVUDRBOTlqdmY4VS1pNzBCVXNlRDLSAcwBQVVfeXFMUEV5OEFOaFpYZlV0bmpGd0NZdzB6N2tBSUg1YUJMNUJoZV8wLTU3RVhuQzRwUm9rRHR5Y2MxNElOZ0pqdFQ4RGlfMTBQeG1OVWd5NFBsRXZHdHFYX0ZqOXZwNWszU2VQVVNJQUFuRERIOC1wOTFKRWlBeWRRUjF5TnhrV0FPWHB0aktmWVExeWxsZnA2dTVrVEwtZlY1RzdKaGh1YnljMmlRZ3ctVVVnTnhIb1djSnlXQU03Vnp4NVl6QlNhb1hWaUhBMGhj?oc=5","published":"Wed, 22 Jul 2026 08:30:12 GMT"},{"title":"[AINews] AI Cybersecurity becomes top of mind","summary":"A roundup notes a cluster of new cybersecurity-focused AI headlines this week, signaling security is becoming a recurring editorial theme rather than a one-off story.","source":"latent_space","url":"https://www.latent.space/p/ainews-ai-cybersecurity-becomes-top","published":"Wed, 22 Jul 2026 03:27:29 GMT"}]},{"name":"Kimi K3: The Model That Reshuffled the Market","slug":"kimi-k3-reshuffled-market","summary":"Moonshot's Kimi K3 forced simultaneous engineering, business, and political responses today: vLLM shipped optimized production serving, Microsoft is weighing a swap into Copilot to cut costs, and the White House escalated IP-theft and chip-smuggling accusations against Moonshot.","articles":[{"title":"A Preview of Production-Scale Kimi K3 Support on vLLM","summary":"vLLM previewed production-scale Kimi K3 serving, including KDA-aware prefix caching, fused kernels, optimized MXFP4 MoE, and multimodal support on both NVIDIA and AMD paths.","source":"vllm_blog","url":"https://vllm.ai/blog/2026-07-22-kimi-k3-preview","published":"Wed, 22 Jul 2026 00:00:00 GMT"},{"title":"Microsoft considers replacing ChatGPT and Claude with China's Kimi K3 to save $600 million - moneywise.com","summary":"Microsoft is reportedly evaluating Kimi K3 as a Copilot backend to cut roughly $600M in inference costs versus its current OpenAI and Anthropic model mix.","source":"search_cn_open_weight_labs","url":"https://news.google.com/rss/articles/CBMijwFBVV95cUxOWlktWHhrNFZCNkVRNjJ1VVYtcTI3bEJ6ZV9QVmtOcUd0Tkl5MExwWVZGMUhOc3NsS0NVLUo3M0VSdmNQOC1JMk9yZ2VGMG0wRmtpTkE2OUpMSl9GdWljekwxVnZlOHpIRlhldkNCYXN1SzZkcU5fNm02OEk0WThTWXdtUnBKTmRQczJNalV5Zw?oc=5","published":"Wed, 22 Jul 2026 09:30:07 GMT"},{"title":"Open models recap: more on Kimi K3, Qwen 3.8, Xi's WAIC speech, distillation, the open-closed gap, and what's next","summary":"Interconnects' recap digs into whether Kimi K3 and Qwen 3.8 were distilled from closed frontier models, and what that implies for how fast the open-closed performance gap keeps closing.","source":"interconnects","url":"https://www.interconnects.ai/p/open-models-recap-more-on-kimi-k3","published":"Wed, 22 Jul 2026 14:09:04 GMT"},{"title":"White House accused China's Moonshot AI of accessing banned Nvidia chips and stealing from Anthropic - qz.com","summary":"A White House official accused Moonshot AI of training Kimi K3 using export-controlled Nvidia chips and technology taken from Anthropic, escalating the dispute beyond generic IP claims.","source":"search_cn_open_weight_labs","url":"https://news.google.com/rss/articles/CBMiggFBVV95cUxQZUhHM25QSk9sTlQwMEN5RGFNT2pTN2VqZ2JCcHBNcWVEVDFxNkJVblE4Wld3eGFJTWpGd1c2ZzNVbXdGR1BKUHFMRnJBN2VlbkI4QmN2M3BWSWxxZi0tNFBoTXVmMVE2QnJFTTJxdmVfNHJOa3ZmUkw0N19FanItMGVn?oc=5","published":"Wed, 22 Jul 2026 17:22:50 GMT"},{"title":"OpenAI President Says Kimi K3 \"Pretty Good,\" Unsure If Distilled - Bloomberg News - TradingView","summary":"OpenAI's president publicly called Kimi K3 \"pretty good\" while declining to confirm whether it was distilled from a closed frontier model, a notably measured reaction from a direct competitor.","source":"search_cn_open_weight_labs","url":"https://news.google.com/rss/articles/CBMi3wFBVV95cUxQWkM5cHBaN0RqMVZ5OHotZWc5anAxOU9fWjEwOW40cmN6ajRhUzZiVVlNanc1S1VwSGxIY1p6cG42Skp3T2xTXzg5R1pIaWVLbVBxS2prM1NQRDZGdmlpMnk1cmk4c1YtMVozX1JNbTNCYkhvQ2IxN3dpd3BWVFJHVzJjdkszOS1pdEZ1Y1FoeWM2RUwtaGVYSjE5eDFReWF6b1VVMVRSQW9teHhUMXpObThsNUNFV3o4STVGMnVGMjRuMGxOVkJDWnltbDNCLTBVd0Z2eWxOZ2wyYlJheFVR?oc=5","published":"Wed, 22 Jul 2026 19:43:22 GMT"}]},{"name":"Coding-Agent Tooling: Show HN Roundup","slug":"coding-agent-tooling-show-hn","summary":"A cluster of new tools targets how builders run and pay for coding agents: a self-hosted router to cut per-call cost, two native-Mac control surfaces for running multiple agents at once, and a fresh breakdown of what Copilot's usage-based billing actually buys versus raw API access.","articles":[{"title":"Show HN: Millwright – Rust-based, self-hosted LLM router","summary":"Millwright is a self-hosted, Rust-based LLM router built for cost control and transparency, a response to hosted routers proliferating (and OpenRouter's possible acquisition) without an open alternative.","source":"hackernews_ai","url":"https://github.com/Northwood-Systems/millwright","published":"Wed, 22 Jul 2026 19:03:50 +0000"},{"title":"Copilot vs. raw API access: What are you actually paying for?","summary":"GitHub breaks down what Copilot's now-listed-API-rate billing buys beyond raw model access: the coding workflow, policy controls, and harness engineering wrapped around the calls.","source":"github_blog_ai_ml","url":"https://github.blog/ai-and-ml/github-copilot/copilot-vs-raw-api-access-what-are-you-actually-paying-for/","published":"Wed, 22 Jul 2026 19:00:00 +0000"},{"title":"Rabbitty – a native Mac terminal for running AI coding agents in parallel","summary":"Rabbitty is a native macOS terminal built specifically to run and manage multiple AI coding agents side by side.","source":"hackernews_ai","url":"https://github.com/mauscoelho/rabbitty-app/releases","published":"Wed, 22 Jul 2026 18:42:28 +0000"},{"title":"Show HN: Forkbench – a native-Mac control room for running CLI coding agent","summary":"Forkbench is another native-Mac control surface for supervising CLI coding agents, part of a fast-forming category of agent-fleet management tools.","source":"hackernews_ai","url":"https://forkbench.com","published":"Wed, 22 Jul 2026 11:31:00 +0000"},{"title":"Show HN: Langy, an automated AI engineer (we gave it a robot body) [video]","summary":"Langy reads a platform's production traces, writes Scenario tests and evaluations for problems it finds, and opens a pull request against the repo on its own.","source":"hackernews_ai","url":"https://langwatch.ai/blog/introducing-langy-your-automated-ai-engineer","published":"Wed, 22 Jul 2026 14:52:14 +0000"}]},{"name":"Platform Bets: Enterprise Agents and Research Funding","slug":"platform-bets-enterprise-agents-research-funding","summary":"The frontier labs made three large forward-looking commitments: OpenAI launched an enterprise voice/chat agent platform, Anthropic funded external economic research on AI's labor impact, and Google earmarked compute credits for scientific-discovery research.","articles":[{"title":"Introducing OpenAI Presence","summary":"OpenAI launched Presence, positioned as a proven enterprise agent platform for deploying trusted voice and chat agents across customer and internal workflows.","source":"openai_blog","url":"https://openai.com/index/introducing-openai-presence","published":"Wed, 22 Jul 2026 05:30:00 GMT"},{"title":"Supporting ambitious external research through the Anthropic Economic Futures Research Fund","summary":"Anthropic committed $200 million to fund external research on AI's economic effects, aiming to get independent data ahead of policy debates rather than after them.","source":"anthropic_newsroom","url":"https://www.anthropic.com/news/economic-futures-research-fund-agenda","published":"2026-07-22T17:00:00+00:00"},{"title":"Accelerating the frontiers of scientific discovery: Google's $40M commitment to the Genesis Mission","summary":"Google committed $40M in AI tokens and compute credits to the Genesis Mission, a push to apply frontier models directly to open scientific-discovery problems.","source":"google_deepmind_blog","url":"https://deepmind.google/blog/accelerating-the-frontiers-of-scientific-discovery-googles-40m-commitment-to-the-genesis-mission/","published":"Wed, 22 Jul 2026 13:38:54 +0000"}]}]}