Story
arxiv_cs_ai ยท Jun 8, 2026 ยท paper
arxiv.orgJun 8, 2026
original source linked
In brief
Vision-language model (VLM) agents are increasingly deployed in interactive game environments. Yet game benchmarks for VLM agents typically report a single first-attempt score per (agent, game) pair, focus on single-a...