Story

arxiv_cs_lg ยท Oct 3, 2026 ยท paper

Source brief

PhaseGate: Phase-Aware CPU Retrieval Scheduling for On-Device LLMs on Unified Memory

arxiv.orgOct 3, 2026
original source linked

In brief

On-device assistants run GPU-based LLM inference alongside CPU retrieval on unified-memory systems. Under a saturated local-retrieval workload, four concurrent retrieval workers raise 95th-percentile (p95) decode late...

Feed lens
eval

Continue reading

Read the original at arxiv.org โ†’Open in live feedRead that dayโ€™s brief

Earlier in this thread 4 items