Story
arxiv_cs_cl ยท Aug 14, 2026 ยท paper
Source brief
Batch-wise Adaptive Pruning: Periodic Neuron Activation-Aware Weight Pruning for Language Reasoning Model
arxiv.orgAug 14, 2026
original source linked
In brief
Large Reasoning Models (LRMs) achieve strong performance on complex tasks through extended chain-of-thought generation, but incur substantial computational costs during inference. In production settings, batched infer...
Feed lens
eval