当熵不够用时:在大语言模型输出长度预测中恢复丢失的语义
When Entropy Is Not Enough: Reclaiming Lost Semantics in LLM Output Length Prediction
AI总结:
该研究针对 LLM 服务中序列填充导致的算力浪费问题,提出 ESTP 框架结合熵与语义重要性预测输出长度,在 ForeLen 基准及端到端测试中均表现更优,可提升吞吐量并降低填充率。
AI中文摘要:
高效的大语言模型(LLM)服务常因需将序列填充至固定最大长度而遭遇瓶颈,这会浪费算力并降低吞吐量。提前预测输出长度可采用感知长度的调度策略,从而减少开销,该优势在长上下文推理与强化学习应用中尤为显著。现有方法如熵引导的 token 池化,以逐 token 熵为主要信号,但往往忽略不同 token 间的语义内容差异,导致重要 token 常被低估,而携带少量信息的 token 却被过度强调,损害了长度预测的可靠性。我们提出 ESTP(熵与语义 token 池化),这是一个轻量级框架,通过结合熵与基于注意力的重要性分数解决该问题;这些分数直接来源于 LLM 预填充阶段计算的自注意力权重,使 ESTP 能以极少的额外计算量同时捕捉不确定性与语义重要性。由于该框架复用预填充激活值,几乎不增加额外内存开销,且仅引入极小延迟。在 ForeLen 基准测试中,ESTP 优于基线方法,在多数场景下实现了更高的预测精度与更低的错误率;在端到端系统测试中与感知长度的调度器集成后,还进一步提升了整体吞吐量并降低了填充率。我们的研究为感知长度的 LLM 服务系统提供了实用且有效的构建模块。
英文摘要:
Efficient LLM serving is often bottlenecked by the need to pad sequences to a fixed maximum length, and this wastes compute and degrades throughput. Predicting output lengths in advance makes it possible to adopt length-aware scheduling, and this reduces the overhead. This advantage is especially pronounced in long-context reasoning and reinforcement learning applications. Existing approaches, such as entropy-guided token pooling, use token-wise entropy as their primary signal, but they tend to ignore differences in semantic content across tokens. So, important tokens are often underweighted, and tokens carrying little information receive disproportionate emphasis. This hurts the reliability of length prediction. We introduce ESTP (Entropy-and-Semantic Token Pooling), a lightweight framework that addresses this issue by combining entropy with attention-based importance scores. These scores are derived directly from the self-attention weights computed during the LLM prefill phase, and this allows ESTP to capture both uncertainty and semantic importance with minimal additional computation. Since the framework reuses prefill activations, it adds almost no extra memory overhead and introduces only minimal latency. On the ForeLen benchmark, ESTP outperforms baseline methods, achieves better prediction accuracy and lower error rates in most scenarios. When integrated with a length-aware scheduler in end-to-end system tests, it further helps improve overall throughput and reduce the padding ratio. Our results offer a practical and effective building block for length-aware LLM serving systems.