Story

langchain_blog ยท Jul 24, 2026 ยท news

Source brief

How We Benchmark Deep Agents

the LangChain blogJul 24, 2026
original source linked

In brief

We revamped how we benchmark Deep Agents. Here's the eval setup we run in Harbor across coding, conversation, and retrieval, and how we use it to ship changes.

Feed lens
agenteval

Continue reading

Read the original at langchain.com โ†’Open in live feedRead that dayโ€™s brief

Earlier in this thread 4 items