{"date":"2026-09-08","title":"What happened in AI — Sep 8, 2026","generated_at":"2026-09-08T21:14:23Z","intro":["DeepSeek's hiring spree is the day's clearest signal: 150 open engineering roles and zero for AI researchers, paired with a plan for a 1-gigawatt data center on 160,000 Huawei chips — a frontier lab betting its bottleneck has shifted from model research to systems and deployment at scale.","Agent security got a reality check too: a new CVE and GitLab's own internal red-team both found agents escaping sandboxes through network access alone, not code execution, while LangChain, Databricks, and a new open-source memory store independently converged on structured, persistent context as the fix for multi-agent sprawl."],"highlights":["DeepSeek is hiring 150 engineers with zero openings for AI researchers — a bet that its bottleneck is now systems and applications, not model quality.","DeepSeek also revealed a 160,000-Huawei-chip, 1-gigawatt data center plan, the infrastructure bet behind that pivot.","A CVE and GitLab's own internal eval both found coding agents escaping their sandboxes through network access alone — isolation without network lockdown isn't isolation.","LangChain, Databricks, and a new open-source memory database all shipped ways to give agents structured, persistent context instead of one shared blob.","vLLM's AgentX optimizations hit up to 130K tokens per GPU-second, a 14.6x-106x serving-cost advantage on SemiAnalysis's agentic benchmark."],"article_count":9,"categories":[{"name":"Agent Context & Memory Engineering","slug":"agent-context-memory-engineering","summary":"Three separate vendors converged on the same fix for multi-agent sprawl today: give agents structured, persistent state instead of one shared context blob.","articles":[{"title":"Organizing Context in a Multi-Agent Harness","summary":"LangChain's deepagents library adds context modes so a subagent can fork the supervisor's full context or start isolated, cutting token cost and cross-talk in multi-agent runs.","source":"langchain_blog","url":"https://www.langchain.com/blog/organizing-context-in-a-multi-agent-harness","published":"Tue, 08 Sep 2026 18:07:23 GMT"},{"title":"Build durable agents with Temporal and Lakebase","summary":"Databricks pairs Temporal's durable execution with Lakebase so an agent like a loan-underwriting workflow can pause mid-task awaiting evidence and resume exactly where it left off.","source":"databricks_blog","url":"https://www.databricks.com/blog/build-durable-agents-temporal-and-lakebase","published":"Tue, 08 Sep 2026 15:46:40 GMT"},{"title":"Show HN: Fraise: a memory database for AI agents","summary":"Fraise stores agent memory as a single-binary temporal graph of facts, topics, and entities, queried with two verbs (remember/recall) instead of a vector database.","source":"hackernews_ai","url":"https://docs.getfraise.dev","published":"Tue, 08 Sep 2026 11:49:20 +0000"}]},{"name":"Agentic Inference Serving at Scale","slug":"agentic-inference-serving-at-scale","summary":"Serving-side benchmarks sharpened today: vLLM published cost numbers for agentic workloads specifically, and AWS compared GPU generations on the same small models.","articles":[{"title":"vLLM x AgentX: Optimizing for Real-World Agentic Serving","summary":"vLLM's KV cache management, parallelism, scheduling, and P/D disaggregation hit up to 130K tokens per GPU-second, a 14.6x-106x serving-cost advantage validated on SemiAnalysis's AgentX benchmark.","source":"vllm_blog","url":"https://vllm.ai/blog/2026-09-08-vllm-agentx","published":"Tue, 08 Sep 2026 00:00:00 GMT"},{"title":"Benchmarking small LLM inference on SageMaker AI: G7 vs G5 and G6","summary":"AWS benchmarked Qwen3-Coder-30B and NVIDIA Nemotron-3-Nano-30B (both MoE models) across G5, G6, G6e, and G7 SageMaker instances, comparing throughput, latency, and cost-per-token.","source":"aws_ml_blog","url":"https://aws.amazon.com/blogs/machine-learning/benchmarking-small-llm-inference-on-sagemaker-ai-g7-vs-g5-and-g6/","published":"Tue, 08 Sep 2026 16:21:57 +0000"}]},{"name":"Agent Sandbox Security","slug":"agent-sandbox-security","summary":"Two unrelated reports landed on the same conclusion: an isolated sandbox is not safe if the agent inside it still has network access.","articles":[{"title":"CVE-2026-82533: DeepSeek Harness Vulnerability Lets AI Agents Escape Their Own Sandbox","summary":"OX Security disclosed CVE-2026-82533, a flaw in DeepSeek's agent harness that lets an AI agent escape its own sandbox entirely, not just its coding container.","source":"search_cn_open_weight_labs","publisher_name":"OX Security","publisher_domain":"ox.security","url":"https://news.google.com/rss/articles/CBMijgFBVV95cUxQT1ptSVFXeUQzc1gtNzhzQmpsbGJSdHd1RndnSm0zMWRkcWV5dHNVZk5pWUFzaUZyQzlPZ25GMHhRemY4cVBUUkFhU1JoeUo4Z0RWVlZVNHhnRV9mYWlBNlJXNjZUcC1Dd2dnQ0tmZVMxTkdHMEhhazFuazBQRTlWeUsxWHZwSU0zZE1UZTNn?oc=5","published":"Tue, 08 Sep 2026 20:37:54 GMT"},{"title":"GitLab Warns That AI Agent Sandboxes Are Only as Secure as Their Network Access","summary":"GitLab's internal security evaluation found that isolating a coding agent in a sandbox doesn't make it safe when the sandbox still has network access — its own test agent escaped using only that access.","source":"infoq_ai_ml","url":"https://www.infoq.com/news/2026/09/gitlab-ai-sandbox-access/?utm_campaign=infoq_content&utm_source=infoq&utm_medium=feed&utm_term=AI%2C+ML+%26+Data+Engineering","published":"Tue, 08 Sep 2026 12:00:00 GMT"}]},{"name":"DeepSeek's Pivot From Research to Systems","slug":"deepseeks-pivot-from-research-to-systems","summary":"DeepSeek's hiring spree and infrastructure plans both point the same direction: the lab is now staffing and building for systems and deployment, not model research.","articles":[{"title":"DeepSeek is hiring 150 engineers, and none of them will touch a model","summary":"DeepSeek is hiring 150 engineers across systems and applications roles, with zero openings for AI researchers, signaling its bottleneck has shifted from model quality to shipping product.","source":"search_cn_open_weight_labs","publisher_name":"The New Stack","publisher_domain":"thenewstack.io","url":"https://news.google.com/rss/articles/CBMiX0FVX3lxTE1SZ3pPOEo1aEJMYzRuaGtvUGFXUG1BVkNiaVFQbGpTZDh5Vm9sQ21nUTlYQzhabXdLUkU0UkxDQ01JWkV5QVdXYmVoQjd1WW1CUTFFNFJLalVVR3dNQm04?oc=5","published":"Tue, 08 Sep 2026 19:02:14 GMT"},{"title":"DeepSeek Plans 160,000 Huawei Chips for a 1-GW Data Center","summary":"DeepSeek plans a 1-gigawatt data center built on 160,000 Huawei chips, the infrastructure bet behind its pivot toward large-scale domestic deployment.","source":"search_cn_open_weight_labs","publisher_name":"BBN Times","publisher_domain":"bbntimes.com","url":"https://news.google.com/rss/articles/CBMimgFBVV95cUxOMUpiSnh3X2U5NHhCajROQjJqb2FPR2NPWm94X0ZaZUhRcFZGR2hVeU85eXkwdE1fSmUyWW8ySlotX0ozUEVfLUZOTjk2bWdGdG1OMG84Zi1nbXVJQjFTT3RKNDZ2TGNDUWUwTld5U3JGWUlmWE5UdG1BTDVlSThnaDdDbWk0SmV4M18zMUNfdUY1cmNteFhwVXJ3?oc=5","published":"Tue, 08 Sep 2026 07:18:12 GMT"}]}]}