arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型能否实现“超线程”?

Can Large Language Models "Hyper-Thread"?

Fei Ding

arXiv 2608.22376首次发表:更新:

发表机构

Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出模型超线程假说,设计三种条件评估并行功能加载等,在 AIME 2025 开发集上获最高准确率,为超线程假说提供初步证据,推动推理扩展视角转变。

AI 中文摘要

大语言模型按顺序生成 token,但能否在生成每个 token 的同时并行执行多项任务?更广泛的注意力分配可能为这种任务并行提供机制。现有推理扩展方法主要依赖更长的生成序列、更多样本或额外验证阶段,而注意力分散常被视为干扰或错误的信号,因此串行生成中的任务并行仍未得到充分探索。我们提出模型超线程假说,并利用在同一问题内共享状态的多个协调任务评估其预测结果。我们设计了三种条件:基线(Baseline)、串行功能调度(Serial Functional Scheduling)和并行功能加载(Concurrent Functional Loading),并使用准确率、输出 token 分布和注意力指标评估它们的收益与成本。在 AIME 2025 开发集上,并行功能加载取得了最高准确率。相对于串行功能调度,其典型输出长度相近,且在大多数问题上更短,同时表现出更大的注意力分散和更高的任务相关覆盖度,尽管存在更重的输出长度尾部。步内并行及其因果机制仍需直接测试。我们的结果表明,更分散的注意力可与更高准确率共存,为超线程假说提供了初步的行为和相关证据。这些发现推动推理扩展的视角从“生成更多 token”转向“让每个生成步骤承载更多任务”,为提升推理性能指明了新方向。

英文摘要

Large language models generate tokens sequentially, but can they execute multiple tasks concurrently while forming each token? Broader attention allocation may provide a mechanism for such task concurrency. Existing approaches to scaling inference primarily rely on longer generations, more samples, or additional verification stages, while attention dispersion is often treated as a signal of interference or error. Task concurrency within serial generation therefore remains underexplored. We propose the Model Hyper-Threading Hypothesis and evaluate its predictions using multiple coordinated tasks that share state within the same problem. We design three conditions (Baseline, Serial Functional Scheduling, and Concurrent Functional Loading) and evaluate their benefits and costs using accuracy, output-token distributions, and attention metrics. On an AIME 2025 development set, Concurrent Functional Loading achieves the highest accuracy. Relative to Serial Functional Scheduling, its typical output length is similar and it is shorter on most problems, while exhibiting greater attention dispersion and higher task-relevant coverage, albeit with a heavier output-length tail. Within-step concurrency and its causal mechanism still require direct tests. Our results show that more dispersed attention can coexist with higher accuracy, providing preliminary behavioral and correlational evidence for the hyper-threading hypothesis. These findings motivate a shift in perspective on inference scaling from "generating more tokens" toward "having each generation step carry more tasks," pointing to a new avenue for improving reasoning performance.

Comments12 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑