Story
infoq_ai_ml ยท Oct 9, 2026 ยท news
Source brief
Android Bench 2 Adds Support for Long-Horizon Tasks, Agentic Evaluation, and Continuous Scoring
infoq.comOct 9, 2026
original source linked
In brief
Google has released Android Bench 2.0, a major update to its benchmark framework for evaluating AI models and agents on Android development tasks. The update introduces long-horizon tasks (LHTs), agent-based evaluatio...

Feed lens
agenticevaluation
Continue reading