LLM Digest
Subscribe

Agent Know-How

Agent engineering · know-how

You can't see why an agent did what it did

🧱 Obstacle·observability·active·23 sources·updated 2026-09-24

When an agent does the wrong thing, the run that produced it is a long, non-deterministic chain of model calls, tool results, and intermediate decisions — and most of that is invisible after the fact. Unlike a stack trace, an agent's "why" is spread across a trajectory you didn't log in enough detail, can't replay deterministically, and can't easily diff against a working run. Debugging an agent is increasingly the job, not a footnote to it.

State of the art

Observability for agents is splitting from generic APM into a trace-first discipline: the unit you capture is the full trajectory (prompts, tool calls, results, retries, sub-agent handoffs), and the work is making that trajectory queryable, diffable, and explainable. Tooling is consolidating around a common trace format and then layering analysis on top — open-source debuggers ingest traces from the emerging standards (Langfuse, Arize/OpenInference, or plain JSONL) and run a model *over the traces themselves* to surface recurring failure patterns rather than make an engineer read every span (HALO). Vendors are pushing the same idea up the stack into managed triage: LangSmith now ships a fleet on-call copilot for alert triage and dedicated voice/trace debugging, treating "read the traces and tell me what's breaking" as an agentic product rather than a dashboard. A second front is monitoring agents you can't fully trace at runtime — offline behavior monitoring evaluates internal agents from logged activity after the fact, which matters when live instrumentation is incomplete or the agent runs where you can't watch it. The hard, still-open part is *evaluating the monitoring itself*: a multi-dataset benchmark for LLM agents in microservice failure diagnosis (AgentOps) exists precisely because "did the agent correctly diagnose the failure" is itself a trajectory-grading problem over multimodal observability data — so agent observability and evaluation are converging, with the trace as the shared substrate.

Instrumentation is also showing up inside the coding-agent product itself, not just in third-party observability tooling: Claude Code now emits workflow.run_id and workflow.name as OpenTelemetry attributes, so a multi-agent workflow run is traceable through the same OTel pipeline a team already operates for the rest of its stack, rather than requiring a bespoke exporter. Enterprise case studies are catching up to the same convergence from the ops side: Schneider Electric built its LLMOps foundations on LangSmith specifically to unify observability, evaluation, and deployment at scale — a real deployment of the "trace as shared substrate" idea, not just a vendor pitch for it.

Trace debugging is also going cross-vendor: LangSmith now positions itself as the debug console for whichever coding agent a developer reaches for — Claude Code, Codex, Cursor, or Copilot — inspecting tool calls, sub-agent handoffs, errors, cost, and retries in one place instead of reading each tool's own logs, treating "which agent produced this trace" as a detail the observability layer should abstract away.

A self-hosted control-plane pattern is emerging alongside the managed vendors above: AWS's Claude Apps Gateway is a stateless container an organization runs itself in front of Claude Code/Desktop, relaying per-request usage metrics to the team's own OpenTelemetry collector (CloudWatch, Prometheus) while enforcing YAML-defined spend caps by org, group, or user — folding telemetry relay and cost policy into one customer-owned layer instead of a vendor dashboard. A managed vendor is now meeting that self-hosted instinct partway: LangSmith's Bring Your Own Cloud option reached general availability on AWS, giving an enterprise team managed observability, evaluation, and deployment while the workload itself stays inside their own VPC — the same "keep it in our network" requirement the Claude Apps Gateway answers by self-hosting, here answered by a vendor deploying its managed product into the customer's cloud instead.

Trace-first observability is also widening to a new modality: LangSmith now traces voice agents built on Pipecat, LiveKit, OpenAI Realtime, and Gemini Live, capturing audio, STT/TTS latency, interruptions, and tool calls in one trace — the same trajectory-capture discipline this page tracks for text-based agent loops, extended to the turn-taking and latency-sensitive failure modes specific to a spoken interface (see agent latency for why voice has a harder real-time floor than text).

A named enterprise deployment backs the trace-plus-LLM-analysis pattern with a production system: Expedia's STAR (built on FastAPI, Datadog, Celery, Redis, and Langfuse) ingests service telemetry during live incidents, runs it through structured workflows to generate root-cause assessments, and keeps engineers in the loop for the final call rather than auto-resolving — an instance of the trace-first, agentic-analysis pattern (HALO, LangSmith's on-call copilot) built on infrastructure a platform team already runs, not a new observability product.

A named experiment sharpens where the RCA bottleneck actually sits: a Coroot test running root-cause analysis across eleven models finds LLMs can already do the reasoning once given correctly prepared context, which reframes the hard problem from "can the model reason about the failure" to "can the pipeline correlate telemetry into that context" — the same context-assembly work Expedia's STAR already invests in rather than a bigger model. The self-hosted, indie tooling layer keeps growing alongside the vendor consolidation this page tracks: a Show HN entrant ships observability specifically for coding agents and LLM applications, one more option in the trace-first tooling space beyond the named vendors above.

A new benchmark puts a number on how far that reasoning-vs-pipeline gap still has to close: ORCA-bench pairs a live, OpenTelemetry-instrumented microservice testbed (six days of metrics, logs, and traces through Prometheus, Jaeger, and OpenSearch) with 1,079 oncall root-cause-analysis tasks graded by an LLM-as-judge independently re-scored by human SREs (agreement κ=0.90). Across five frontier agents the best RCA accuracy is 25.3% on realistic-input tasks and 10.0% on hard ones — a gap that holds even for Claude Fable 5, and the weakest model hallucinates an implausible root cause on 40% of reports. Since the testbed is a curated 50GB slice of a public system, the authors read this as a lower bound on the real-world gap, sharpening the Coroot finding above: the reasoning may already be there, but the end-to-end oncall pipeline this page tracks (telemetry correlation, ambiguous reports, time pressure) is still mostly unsolved.

A named production deployment pairs tracing with a human-approval gate rather than auto-resolving: LangChain built an autonomous SRE agent for Kubernetes on Deep Agents that requires human approval before it applies a change, with every step, tool call, and decision captured in LangSmith traces — an instance of the trace-first, agentic-analysis pattern above (Expedia's STAR, HALO) where the trace is also what a human reviews before the agent is allowed to act, not just what an engineer replays afterward.

A second named deployment pairs the trace-first pattern with production security-ops rather than SRE: Figma built agents on a Panther SIEM foundation, querying over 100 data sources (AWS, Okta, GitHub, GCP, osquery) with an alert-triage agent that reasons over the full Slack thread plus its own steering memory, scoped to the tools an on-call engineer would actually use. The team reports memory — not model choice or tool count — as the lever with the most impact on quality, backed by measured results: 70% faster resolution on complex alerts, a 20% cut in on-call pages from re-tuned severity, and 100+ previously-unknown vulnerabilities surfaced. Guardrails mirror LangChain's Kubernetes SRE pattern above rather than trusting the agent's own judgment: agent-authored PRs default to draft status and every fix still needs human approval before it ships.

Capture tooling itself is widening on the open-source side: Simon Willison's llm CLI (0.32) adds support for visible reasoning traces and redesigned, smarter logging alongside server-side provider tools — the same trajectory-capture discipline the vendor platforms above ship, now available in a widely-used, framework-agnostic command-line tool rather than only a hosted product.

A major serving platform now ships tracing natively rather than leaving it to a third-party SDK: Cloudflare added agent tracing directly into existing Workers traces, with invoke_agent → chat/execute_tool → tool_approval spans keyed by agent name, agent ID, and conversation ID so a session replays turn by turn. The launch also exposes the privacy tension this page's trace-first shift creates rather than solves: message and tool payloads default to *not* being stored under one SDK wrapper (Vercel's AI SDK) but *are* stored by default under another (Flue) — the same platform feature ships with opposite privacy defaults depending on which harness a team already picked, and those payloads routinely carry personal data or secrets. Payloads are also subject to undisclosed span-size truncation, so a trace can silently drop the reasoning or tool arguments a debugging session needed most.

The trace-first thesis has a boundary condition when agents talk to each other: work on Verifiable Latent Alignments (VLA) starts from the fact that language-model agents can coordinate through continuous hidden states that never appear in the public transcript, so a complete trace of what was *said* can still miss what was *communicated*. VLA links each private latent-state record and channel status to the resulting public action through a shared event identifier, so a monitor can causally match a decision against the hidden channel that produced it, and combines representation anomaly detection into a layered monitor rather than reading transcripts alone. It sharpens what "capture the full trajectory" has to mean in a multi-agent system (see multi-agent): the span schema this page tracks records messages and tool calls, and that is the wrong unit when the coordination happens below the message layer.

The trace format is also gaining a visual payload, not just structured spans: Amazon OpenSearch Service's MCP Apps return an interactive visualization alongside every tool call's text response, rendered inline in the IDE conversation instead of a separate dashboard. A local MCP server authenticates with AWS credentials, forwards the agent's query to the same OpenSearch/Prometheus data sources that power existing dashboards, and returns both a text summary and a rendered widget — letting one conversation move from alert triage to log clustering to trace waterfalls to service-dependency maps without the engineer leaving the chat to open a separate observability tool. It's MCP carrying the observability payload itself, not just the query that produces it.

Session traces and cost controls are converging into one diagnostic signal rather than two separate dashboards: industry coverage of agent observability practice names spotting tool-call loops and runaway spend as the same triage step, since both symptoms show up in the same trace and both need enough preserved execution context for post-incident debugging — the same cost/observability convergence this page's Claude Apps Gateway coverage already tracks, now framed as a general diagnostic pattern rather than one vendor's product.

The trace viewer itself is now catching up to the trace-first thesis this page tracks: LangSmith's Trajectories renders a full agent session as one chat-style thread — user turns, tool calls, and sub-agent handoffs read top to bottom — instead of the tree of nested spans a generic OpenTelemetry viewer shows, trading completeness of the span tree for the specific job this page keeps naming as the bottleneck: an engineer scanning a long-running session fast enough to spot where it went wrong.

What's new

LangSmith shipped Trajectories, a chat-style read of an entire agent session — user turns, tool calls, and sub-agent handoffs in one scrollable thread — instead of the tree of nested spans a generic trace viewer shows, aimed at making a long-running session fast to scan without expanding every span by hand (see State of the art above).

Prior update: Session traces and cost controls are converging into one agent-failure diagnostic: industry coverage frames spotting tool-call loops and runaway spend as the same triage step, both read off the same preserved trace (see State of the art above).

Prior update: Figma built a named production deployment pairing the trace-first pattern with security-ops: an alert-triage agent on a Panther SIEM foundation, scoped to on-call tools and querying 100+ data sources, reports memory as the biggest quality lever and posts measured results (70% faster resolution, 20% fewer pages, 100+ vulnerabilities found) while keeping agent-authored PRs in draft pending human approval (see State of the art above).

Prior update: Amazon OpenSearch Service's MCP Apps return interactive visualizations (trace waterfalls, service maps, log clusters) inline in an agent conversation instead of a text summary alone, letting an investigation move from alert to root cause in one thread instead of switching to a separate dashboard (see State of the art above).

Prior update: Latent-channel monitoring marks the first real limit on this page's trace-first stance: agents can coordinate through hidden states invisible in the public transcript, so message-and-tool-call spans are not a complete record of a multi-agent run. VLA's answer is to link each private latent record to the public action it caused via a shared event identifier — a monitoring unit below the span, not a better span.

Prior update: Cloudflare shipped agent tracing natively into its existing Workers traces, but the launch also surfaces a real gotcha: message/tool payload storage defaults are opposite between its two supported SDKs (off by default in one, on by default in the other), so the same platform feature can silently retain or silently drop personal data depending on which harness a team already chose.

Why it matters for platform engineers

You cannot operate what you cannot explain. Without trajectory-level traces, a regression after a model upgrade, a silent tool failure, or a runaway loop is invisible until it shows up as cost or a user complaint — and you have no way to reproduce it. Observability is the precondition for the rest of the stack: evaluation needs traces to grade, cost control needs per-step attribution, and incident response needs a replayable run. The build-vs-buy question is whether to standardize on a trace format and own the analysis, or adopt a managed platform — but either way the trace is the new log line.

Evidence · 23 sources