Story

langchain_blog ยท Jul 23, 2026 ยท news

Source brief

How We Benchmark Deep Agents

the LangChain blogJul 23, 2026
original source linked

In brief

We revamped how we benchmark Deep Agents. Here's the eval setup we run in Harbor across coding, conversation, and retrieval, and how we use it to ship changes.

Feed lens
agenteval

Continue reading

Read the original at langchain.com โ†’Open in live feed

Earlier in this thread 4 items