arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FLEET:从对数几率熵到文本生成中的增强轨迹

FLEET: From Logits Entropy to Enhanced Trajectories in Text Generation

Oleksii Streltsov, Oleksandra Vitko

arXiv 2609.27657首次发表:更新:

发表机构

Kharkiv National University of Radio Electronics(哈尔科夫国立无线电电子大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FLEET通过记忆机制将生成表示为熵阈值之上的稀疏轨迹,推断逐标记效用分数调整对数几率,在保持相同准确率下实现3倍加速,并将LiveCodeBench Pass@32从59.9%提升至66.2%。

AI 中文摘要

基于大型语言模型(LLM)的解决方案通常依赖温度采样,通过聚合完成分布中的多个样本来提高准确性和稳定性。然而,这种无记忆的方法本质上存在次优性:由于缺乏对先前生成及其评估的感知,随着抽取更多样本,它会产生越来越多的语义重复答案,导致收益递减。为解决这一局限,我们提出了FLEET,一种将记忆机制集成到生成过程中的新颖方法。FLEET将每次生成表示为一条稀疏轨迹,穿过熵超过预定义阈值的状态,并利用这些轨迹推断每个标记的效用分数,从而调整对数几率。基准评估表明,FLEET在达到与重复采样基线相同准确性的同时,实现了3倍加速,并在相同预算下显著提高了复杂编码任务的准确性(LiveCodeBench Pass@32从59.9%提升至66.2%)。此外,在本文评估的贪心解码配置中,该方法具有确定性,并通过单次校准过程推导其主要超参数,仅需对现有LLM流水线进行最小修改。

英文摘要

Solutions based on large language models (LLMs) often rely on temperature sampling to improve accuracy and stability by aggregating multiple samples from the completion distribution. However, this memoryless approach is inherently suboptimal: because it lacks awareness of prior generations and their evaluations, it produces an increasing proportion of semantically duplicate answers as more samples are drawn, leading to diminishing returns. To address this limitation, we introduce FLEET, a novel method that integrates a memory mechanism into the generation process. FLEET represents each generation as a sparse trajectory through states whose entropy exceeds a predefined threshold and uses these trajectories to infer per-token utility scores that adjust the logits. Benchmark evaluations demonstrate that FLEET achieves the same accuracy as the repeated sampling baseline, with a 3x speedup, and substantially improves accuracy on complex coding tasks (LiveCodeBench Pass@32 increases from 59.9% to 66.2%) under the same budget. Furthermore, in the greedy-decoding configuration evaluated here, the approach is deterministic and uses a single calibration pass to derive its principal hyperparameters, requiring only minimal modifications to existing LLM pipelines.

Comments25 pages, 8 figures. Algorithm source code and experiments: https://github.com/Alexiush/fleet

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑