arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估用于路由最终层注意力的轨迹特征

Evaluating Trajectory Features for Routing Final-Layer Attention

Yupeng Yao

arXiv 2610.09272首次发表:更新:

AI 中文总结

本研究评估轨迹特征对最终层注意力路由的增量价值,发现其无显著增益,并指出分配质量与推理收益的分离。

AI 中文摘要

注意力路由需要一种信号,用于预测当前前缀上注意力的价值。我们评估了隐藏状态外推误差、曲率和误差变化是否能在不确定性、一步位移、位置和状态投影的基础上改进这一预测。在冻结的SmolLM3-3B-Base和Qwen3.5-4B-Base检查点上,对最终注意力层的成对执行提供了有符号的下一词元损失差异。在相同的因果20%调用配额下,对100本保留的PG-19书籍测试了效用监督的路由器。在族系校正后,六个预先指定的比较中没有一个显示出正向增益。在Qwen3.5中,参数匹配的固定投影控制相对于轨迹路由器将NLL降低了0.00356 nats/词元(95%区间为0.00218至0.00487)。次要结果取决于被移除的操作、特征位置和评分范围;冻结阈值在更长的范围内也大幅漂移。实际选定的查询执行在增加NLL的情况下带来了较小的长序列延迟降低,而学习到的路由器在缓存续写期间仍然较慢。该研究确定了这些轨迹摘要的增量价值的局限性,并将分配质量与实测推理收益区分开来。

英文摘要

Attention routing requires a signal that predicts the value of attention on the current prefix. We evaluate whether hidden-state extrapolation error, curvature and error change improve this prediction beyond uncertainty, one-step displacement, position and state projections. Paired executions of the final attention layer supply signed next-token loss differences in frozen SmolLM3-3B-Base and Qwen3.5-4B-Base checkpoints. Utility-supervised routers are tested on 100 held-out PG-19 books at an identical causal 20 percent invocation quota. None of six prespecified comparisons shows a positive gain after familywise correction. In Qwen3.5, a parameter-matched fixed-projection control lowers NLL by 0.00356 nats/token relative to the trajectory router (95 percent interval 0.00218 to 0.00487). Secondary results depend on the operation removed, feature location and scoring horizon; frozen thresholds also drift substantially at longer horizons. Actual selected-query execution yields small long-sequence latency reductions with increased NLL, while learned routers remain slower during cached continuation. The study identifies limits on the incremental value of these trajectory summaries and separates allocation quality from measured inference benefit.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑