Story

arxiv_cs_cl ยท Aug 4, 2026 ยท paper

Source brief

SocietyBench: Forecasting Counterfactual Social-World Evolution

arxiv.orgAug 4, 2026
original source linked

In brief

Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task -- fix a bug, drive a browser, operate a GUI. A complementary social ability, namely how...

Feed lens
agenteval

Continue reading

Read the original at arxiv.org โ†’Open in live feedRead that dayโ€™s brief

Earlier in this thread 4 items