Story
arxiv_cs_lg ยท Oct 3, 2026 ยท paper
arxiv.orgOct 3, 2026
original source linked
In brief
On-device assistants run GPU-based LLM inference alongside CPU retrieval on unified-memory systems. Under a saturated local-retrieval workload, four concurrent retrieval workers raise 95th-percentile (p95) decode late...
Feed lens
eval