Story
arxiv_cs_lg ยท Jun 17, 2026 ยท paper
arxiv.orgJun 17, 2026
original source linked
In brief
Embodied Vision-Language-Action (VLA) models are typically obtained by fine-tuning powerful pretrained VLMs on robotics data, yet it is unclear how much commonsense and factual knowledge they retain after adaptation....
Feed lens
agentevaluation