As agent systems scale to consumer-facing products, practitioners are surfacing the operational seams — architecture complexity, model-version regressions, and eval rigor — that pilots don't expose.
An external reconstruction maps how Memory, Proactivity, Scheduling, Browser Use, Plugins, Skills, and Tools combine into ChatGPT Work's agent architecture.
Steve Yegge says his reusable agent framework Gas Town worked brilliantly through Claude Opus 4.6, then fell apart at the seams once Opus 4.7 changed agent behavior.
LangSmith's new guide evaluates voice agents across execution, outcomes, and caller experience using traces, code evaluators, and LLM judges instead of transcript spot-checks.