Story
arxiv_cs_cl ยท Jun 5, 2026 ยท paper
arxiv.orgJun 5, 2026
original source linked
In brief
Current Vision-Language Models struggle with hours-long videos because processing full-length visual sequences induces prohibitive token explosion and attention dilution. To overcome this, we introduce MemDreamer to d...