arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AutoCRAT:针对大语言模型推理的轨迹内随机性与计算量联合控制

AutoCRAT: Within-trajectory Joint Control of Stochasticity and Compute for LLM Reasoning

Hanjun Luo, Qiushi Liu, Jingya Zhang, Haihong Pang, Jiaheng Wen, Yifei Ma, Yu Yao, Chengxi Zhang, Hanrong Zhang, Yankai Chen, Hanan Salam

arXiv 2608.29988首次发表:更新:

发表机构

New York University; New York University Abu Dhabi; University of Washington Seattle; Harvard University; Massachusetts Institute of Technology; University of Illinois Chicago; Mohamed bin Zayed University of Artificial Intelligence(纽约大学; 纽约大学阿布扎比分校; 华盛顿大学西雅图分校; 哈佛大学; 麻省理工学院; 伊利诺伊大学芝加哥分校; 穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AutoCRAT是针对大语言模型推理的轨迹内随机性与计算量联合控制方法,仅在语义边界更新决策,在6个基准上减少推理令牌并提升准确率,且跨主干迁移性强。

AI 中文摘要

大语言模型(LLM)的强推理性能高度依赖推理时的决策,但这些决策通常由静态、一刀切的策略处理,限制了对不同任务和推理阶段的适应性。近期的自适应方法部分解决了这一局限,但它们主要单独适配解码随机性(模型的探索方式)或推理计算量(模型的推理时长),未建模单次推理轨迹内两者的交互。为应对这一挑战,本文转向轨迹内联合控制视角,并将其实例化为AutoCRAT——一种针对冻结主干模型的解码器侧控制器。AutoCRAT仅利用解码阶段可用的信号,在生成过程中联合调整采样随机性与推理预算;其在离散动作空间上运行,仅在语义边界更新控制决策,提升稳定性的同时对不断演进的推理过程保持响应性。在6个基准测试上的全面评估显示,AutoCRAT:(I)与推荐的静态配置相比,平均使用的推理令牌减少13.8%-52.7%;(II)在相对准确率上超出推荐的静态基线和自适应基线1.5%-4.5%;(III)具备强跨主干迁移能力。

英文摘要

Large language models (LLMs) achieve strong reasoning performance, which depends critically on inference-time decisions. Yet these decisions are commonly handled by static, one-size-fits-all policies, limiting adaptation to diverse tasks and reasoning stages. Recent adaptive methods partially address this limitation, but they primarily adapt either decoding stochasticity (how the model explores) or reasoning compute (how long the model reasons) in isolation, leaving their interaction within a single reasoning trajectory unmodeled. To address this challenge, we shift toward a within-trajectory joint control view, and instantiate it in AutoCRAT, a decoder-side controller for frozen backbones. Using only signals available during decoding, AutoCRAT jointly adjusts sampling stochasticity and reasoning budget during generation. AutoCRAT operates over a discrete action space and updates control decisions only at semantic boundaries, improving stability while remaining responsive to the evolving reasoning process. Comprehensive evaluation across 6 benchmarks demonstrates that AutoCRAT (I) uses 13.8-52.7% fewer inference tokens on average than recommended static configurations, (II) surpasses recommended static and adaptive baselines by 1.5-4.5% in relative accuracy, and (III) enjoys strong cross-backbone transferability.

CommentsEMNLP 2026 Findings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑