{"date":"2026-07-23","title":"What happened in AI — Jul 23, 2026","generated_at":"2026-07-23T21:35:00Z","intro":["AWS and LangChain each published production eval blueprints today: Motorway's pipeline on Strands Agents and AgentCore cut wrong results from 1 in 8 queries to 1 in 50 and cut incident-detection time from hours to minutes, while LangChain detailed the Harbor-based benchmark it runs across coding, conversation, and retrieval before every Deep Agents change ships.","The White House also escalated its case that Moonshot AI distilled Anthropic's Fable model to build Kimi K3, a claim Beijing rejected as evidence of Chinese \"self-reliance\" — even as Microsoft confirmed it is still evaluating K3 for Copilot and Nvidia's CEO downplayed the competitive threat."],"highlights":["AWS's Motorway pipeline (Strands Agents + AgentCore) cut wrong agent results from 1 in 8 queries to 1 in 50, and cut incident-detection time from hours to minutes.","LangChain detailed the Harbor-based benchmark it runs across coding, conversation, and retrieval before every Deep Agents change ships.","The White House escalated its claim that Moonshot AI distilled Anthropic's Fable model to build Kimi K3; China called its progress \"self-reliance.\"","Microsoft is still evaluating Kimi K3 for Copilot despite the dispute, and Nvidia's CEO downplayed the competitive threat from Chinese open models.","vLLM shipped an AFD plugin disaggregating attention and FFN for MoE serving across GPU and Ascend NPU backends."],"article_count":19,"categories":[{"name":"Agent Engineering: Evals Get Production Numbers","slug":"agent-engineering-evals-get-production-numbers","summary":"Two vendors moved eval rigor from talk to measured production outcomes today, and two more benchmarks target the risk side of agent evaluation.","articles":[{"title":"Evaluating AI Agents: A production blueprint with Strands and AgentCore","summary":"Motorway and AWS built an eval pipeline on Strands Agents and AgentCore that cut wrong results from 1 in 8 queries to 1 in 50 and slashed incident-detection time from hours to minutes.","source":"aws_ml_blog","url":"https://aws.amazon.com/blogs/machine-learning/evaluating-ai-agents-a-production-blueprint-with-strands-and-agentcore/","published":"Thu, 23 Jul 2026 17:00:20 +0000"},{"title":"How We Benchmark Deep Agents","summary":"LangChain runs its Deep Agents benchmark in Harbor across coding, conversation, and retrieval, and uses it to validate every shipped change.","source":"langchain_blog","url":"https://www.langchain.com/blog/how-we-benchmark-deep-agents","published":"Thu, 23 Jul 2026 17:55:37 GMT"},{"title":"Expedia Uses AI Driven Service Telemetry Analyzer to Accelerate Incident Investigation","summary":"Expedia's STAR platform pairs service telemetry with LLMs, built on FastAPI, Datadog, Celery, and Redis, to speed production incident triage.","source":"infoq_ai_ml","url":"https://www.infoq.com/news/2026/07/expedia-ai-observability-star/?utm_campaign=infoq_content&utm_source=infoq&utm_medium=feed&utm_term=AI%2C+ML+%26+Data+Engineering","published":"Thu, 23 Jul 2026 14:15:00 GMT"},{"title":"A value-poisoning benchmark for consequential agent actions","summary":"A new benchmark tests whether agents can be manipulated into high-stakes harmful actions through subtly poisoned values or context.","source":"hackernews_ai","url":"https://actionrail.ai/research/value-poisoning-benchmark/","published":"Thu, 23 Jul 2026 16:23:50 +0000"},{"title":"Article: Multi-Agent AI for Production Security Operations: An A2A and MCP Architecture in a 5G Core","summary":"A production A2A/MCP multi-agent architecture inside a 5G core keeps detection rules aligned with a threat landscape that evolves faster than analysts can write rules by hand.","source":"infoq_ai_ml","url":"https://www.infoq.com/articles/multi-agent-security-operations/?utm_campaign=infoq_content&utm_source=infoq&utm_medium=feed&utm_term=AI%2C+ML+%26+Data+Engineering","published":"Thu, 23 Jul 2026 09:00:00 GMT"}]},{"name":"Agent Tooling: New Harnesses and Frameworks Ship","slug":"agent-tooling-new-harnesses-and-frameworks-ship","summary":"Independent builders keep shipping agent harnesses and dev tooling faster than any platform standard has emerged to absorb them.","articles":[{"title":"July 2026: LangChain Newsletter — NemoClaw Blueprint, OpenWiki Brains, and More","summary":"July's roundup covers the NemoClaw Deep Agents blueprint, a free LangSmith Sandboxes trial, a Fleet Slack integration, and reinforcement-learned-memory support in Deep Agents.","source":"langchain_blog","url":"https://www.langchain.com/blog/july-2026-langchain-newsletter","published":"Thu, 23 Jul 2026 18:51:32 GMT"},{"title":"Show HN: Rendi, an agent harness on Trigger.dev without spinning up a VM","summary":"Rendi runs background jobs, schedules, a real browser, and email for agents on Trigger.dev, skipping the usual step of provisioning a VM.","source":"hackernews_ai","url":"https://github.com/mcheemaa/rendi","published":"Thu, 23 Jul 2026 13:44:47 +0000"},{"title":"Show HN: Hanesu – An experimental workflow layer for AI coding agents","summary":"Hanesu adds a workflow layer meant to stop long-running coding agents from losing context or repeating steps as tasks grow.","source":"hackernews_ai","url":"https://github.com/jezmn/hanesu","published":"Thu, 23 Jul 2026 06:58:39 +0000"},{"title":"7 Best Claude Code Alternatives for CLI Agentic Coding - KDnuggets","summary":"A roundup of CLI coding-agent alternatives to Claude Code, useful as a quick landscape check on where the category stands.","source":"search_agent_engineering_news","url":"https://news.google.com/rss/articles/CBMihwFBVV95cUxQamlqYkIwRktfeHdET25sR0xKRHYyQ0NXZXYwT3FfODhTaFRjLXhEU21lNFlCUHdJMm1wTEZEdFhyZ0dJdVJSa2dEc1ZyZGZ0TmlSUFRwSFpZX2xIbkw4STh0MEJrMkRYNVRJMms1d0tQelhrMWFDYTlqM2Z0anRuM1NMYmFZamM?oc=5","published":"Thu, 23 Jul 2026 12:08:57 GMT"},{"title":"Show HN: OpenTrust – Browser trust signals for the AI era","summary":"An open-source TypeScript SDK collects privacy-preserving browser trust signals — automation and virtual-camera detection, liveness — without making the block/allow call itself.","source":"hackernews_ai","url":"https://github.com/rafaelEt/opentrust","published":"Thu, 23 Jul 2026 20:37:47 +0000"}]},{"name":"Serving Infrastructure and Narrower Model Bets","slug":"serving-infrastructure-and-narrower-model-bets","summary":"Labs and infra vendors each shipped a narrower, more specialized bet today: MoE-optimized serving, dedicated defense compute, a compact model beating a far larger open-weights rival, and consumer health-data integration.","articles":[{"title":"Announcing vLLM AFD Plugin: Disaggregating Attention and FFN for Flexible MoE Serving","summary":"The new vLLM plugin disaggregates attention and FFN computation for MoE serving, adding GPU and Ascend NPU backends with connector-based execution.","source":"vllm_blog","url":"https://vllm.ai/blog/2026-07-23-vllm-afd-plugin","published":"Thu, 23 Jul 2026 00:00:00 GMT"},{"title":"NVIDIA AI Supercomputer Comes Online at Naval Postgraduate School","summary":"Jensen Huang commissioned a DGX GB300 system at the Naval Postgraduate School, bringing a full-scale AI research platform online for defense-related work.","source":"nvidia_blog","url":"https://blogs.nvidia.com/blog/naval-postgraduate-school-dgx-ai-supercomputer/","published":"Thu, 23 Jul 2026 02:00:46 +0000"},{"title":"Inside the Model Factory — Eiso Kant, Poolside AI","summary":"Poolside's co-CEO describes the small-team model factory behind Laguna S, a 118B-parameter MoE the company says beats a roughly 1T-parameter open-weights rival.","source":"latent_space","url":"https://www.latent.space/p/poolside","published":"Thu, 23 Jul 2026 05:09:14 GMT"},{"title":"Launching Health in ChatGPT","summary":"Eligible US users can now connect medical records and Apple Health to ChatGPT for more personalized health insights.","source":"openai_blog","url":"https://openai.com/index/health-in-chatgpt","published":"Thu, 23 Jul 2026 00:00:00 GMT"}]},{"name":"Kimi K3: The Theft Claim Hardens While Business Keeps Moving","slug":"kimi-k3-theft-claim-hardens-while-business-keeps-moving","summary":"Washington's case that Moonshot AI distilled Kimi K3 from Anthropic's Fable model escalated today, but neither Beijing's rebuttal nor the dispute itself has slowed enterprise interest in the model.","articles":[{"title":"China's Moonshot AI stole from Anthropic, Trump tech advisor says - BBC","summary":"A White House tech advisor said China's Moonshot AI distilled Anthropic's Fable model to build Kimi K3, the most direct accusation yet in the dispute.","source":"search_cn_open_weight_labs","url":"https://news.google.com/rss/articles/CBMiWkFVX3lxTE1iWFdWWk1rRjktUk5wZnBrNHBnd0xhU01XRmo1dTV0aHpGWmRDcTJhWndnUzV3R25aT2VjSHZxaWNzUUhWZF9MeVZnSkRZZUlJYXJrc3RBM2Vvdw?oc=5","published":"Thu, 23 Jul 2026 00:52:25 GMT"},{"title":"China says AI development comes from ‘greater self-reliance and strength’ amid stolen tech claims - The Hill","summary":"Beijing rejected the theft allegations, framing Kimi K3 and other Chinese models as products of domestic self-reliance rather than copied US technology.","source":"search_cn_open_weight_labs","url":"https://news.google.com/rss/articles/CBMigAFBVV95cUxNSVg4NmdPb1ZJSFE3T3l0bEpGd21JMEJ1M0hIc0NOSW5HWG85Zklpb29OenhTR1FmS2hqS2YxeWotRFplWkJyNlRENjQ4NHJ4NWQ0Zm9QZVlUTmZlUlc2UlVYWmpPTnkyQ2dIaGx1ZjV4RkMxNXVXQmJDZHh0aWVxeNIBhgFBVV95cUxQSHFkSTdvWlFvMHF5MmIweHBSQmJxWlBCRi00R1BRNlptTHFkMTBvbEctN0RfeXlIZVc2X1J5dWsyNVJDYk50Zm1GMU5TeWZfWHVheUd4QTF0bXdDT0ttY3RueFQ4bnV5V2lVSXVQNXg4aEpKMkZTY09hVjEwckhveGJKU3VXUQ?oc=5","published":"Thu, 23 Jul 2026 16:43:00 GMT"},{"title":"Microsoft reportedly evaluates Moonshot AI’s Kimi K3 for Copilot - technode.com","summary":"Despite the ongoing dispute, Microsoft is reportedly still evaluating Kimi K3 for use in Copilot, a sign enterprises aren't waiting for the allegations to resolve.","source":"search_cn_open_weight_labs","url":"https://news.google.com/rss/articles/CBMinAFBVV95cUxQRGJ3Y0RDeWhIaHdIRjlBdjNidW5zd0tBQ0Ffd3dSWGJUamRmUlNWbkRBaXJydHQzb0owb0pKNGE3dWRITV9DOC05Ym1RbzBxMHlTSHlHLWVOMHhrd3ZIUThUSTVnRkdBTU9wNWpqbGFhQnhoRXBWdHN4RW80bjd2ZUlabDZZeE5oZXR2M0RPYUZfZ0pjeGRrTGQ3ZEs?oc=5","published":"Thu, 23 Jul 2026 09:07:35 GMT"},{"title":"Nvidia CEO says the US has nothing to fear from China’s open-source AI models - The American Bazaar","summary":"Nvidia's CEO downplayed the competitive threat from Chinese open-source models, pushing back on the framing driving the Kimi K3 controversy.","source":"search_cn_open_weight_labs","url":"https://news.google.com/rss/articles/CBMitwFBVV95cUxNLVFvSURGWFl6ek5sUDZZaExmRWNxVGJjbTI5RE8wekxFc3ZOejJrLVBuOW5KN08yMFVYc3pOcDdnUzRCVVVadHZlZnpYNS1iUlZhQ2ZmODM5NFNkcDB0LTViZ194WHdUcnYzZzZwVXNmSS0tVzJhM25CSlItcUpaZWwxcTdXazFTRGhqR1VyTThsWEZlRHJtQUVBRXpZWG5LVGVfZEVUMTN3Mm10a0E0LTFmZUVDd3M?oc=5","published":"Thu, 23 Jul 2026 16:52:54 GMT"},{"title":"Moonshot AI model challenges EDA moat with 2-day chip design - digitimes","summary":"Separately from the origin dispute, a Moonshot model reportedly completed a chip-design task in two days, a concrete capability claim against incumbent EDA tools.","source":"search_cn_open_weight_labs","url":"https://news.google.com/rss/articles/CBMiowFBVV95cUxOb2gzSXFpeGVzUnpXVzRSQTlNWGUyNjhmOTdxZ0Z6Sm5KbS1WTTRYS245ZEtYeDlmWUZ5ZjBJeEhIdHF6OWlyVTk5VGdSZjJGQXR1Ni1JOTBZSVprbWNZa25faVdoVm95WWlDdmEzT2hGYlpVVlo0cTB4bVNaYmdWSkJWdnVucVYxVXA4aEdHdTJXZTVaTzBRbUxvOE0yWFZqVDRr?oc=5","published":"Thu, 23 Jul 2026 07:11:54 GMT"}]}]}