Show HN: Boffin – Staff-engineer layer for AI coding agents
Boffin adds a 'staff engineer' layer that routes per-edit architectural constraints into AI coding agents' output before it lands.
13 articles · 3 categories
The finishable daily brief
Sunday, Jul 26, 2026
13 articles · 3 categories
read top to bottom · then stop
In 30 seconds
Two Show HN launches target coding agents' weakest spots today: architecture drift and lost session memory. A third open-source project, OpenLake, claims a 50% cut in long-horizon inference costs by offloading KV caches from GPU memory to RAM/NVMe, while a separate hypervisor project chases the same cost problem from the consumer-compute angle. Simon Willison surfaced the flip side of that cost pressure: an investigation into the relay market reselling pooled LLM API keys, some of it tied to fraud.
The day's bigger industry story is China's open-weight AI push hitting turbulence at the same time it gains ground. Moonshot AI shipped Kimi K3 and a widely syndicated wire report tracked growing US enterprise interest in cheaper Chinese models, even as DeepSeek paused a new funding round days after founder Liang Wenfeng's leaked remarks went viral.
Two new tools attack coding-agent reliability from different angles: Boffin injects per-edit architectural constraints, while CMEM gives agents memory that survives across sessions instead of resetting each run.
Boffin adds a 'staff engineer' layer that routes per-edit architectural constraints into AI coding agents' output before it lands.
CMEM gives coding agents persistent memory across sessions, targeting the context-loss problem that resets an agent's progress between runs.
Two open-source projects chase cheaper inference through KV-cache offloading and consumer-compute hosting, while an investigative report maps a parallel relay market reselling pooled API tokens at a discount.
OpenLake is an open-source storage engine that offloads LLM KV caches from GPU memory to a shared tier of RAM and NVMe, cutting long-horizon inference costs by half.
Scalattice built a hypervisor for running inference workloads on consumer-grade compute, aiming to make idle consumer hardware usable for LLM serving.
Investigative reporting picked up by Simon Willison maps a relay market where pooled API keys resell discounted LLM tokens, some of it tied to fraud.
Moonshot shipped Kimi K3 and a widely syndicated wire report tracked rising US enterprise interest in cheaper Chinese models, while DeepSeek paused a new funding round days after founder Liang Wenfeng's leaked remarks went viral.
Moonshot AI released Kimi K3, drawing renewed US commentary on the pace of Chinese open-weight model releases.
DeepSeek paused a new funding round in the days following founder Liang Wenfeng's viral remarks, per The Times of India.
Leaked remarks attributed to DeepSeek founder Liang Wenfeng went viral, per South China Morning Post, without official confirmation from the company.
A widely syndicated wire report tracks growing US enterprise interest in cheaper, open Chinese models even as export controls remain in place.
You are caught up for this edition