arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03961cs.AI

面向大语言模型测试时缩放的可解释自适应采样

Interpretable Adaptive Sampling for LLM Test-Time Scaling

Mobina Kashaniyan, Ali Jannesari

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对LLM测试时推理的固定计算预算不灵活、不可解释的问题,提出带轻量级模糊控制器的自适应采样方法,在问答和数学任务上提升性能并减少平均采样数,为高效测试时推理提供实用方向。

中文摘要 AI 辅助

测试时缩放通过生成并聚合多个候选答案来提升大语言模型(LLM)的推理能力,但许多流水线采用固定的单查询预算,对简单和困难的提示词投入相同的计算资源。这些固定预算也难以解释,因为它们无法说明某个特定提示词为何获得特定数量的采样。我们提出带有轻量级模糊控制器的自适应测试时缩放方法,该控制器将可解释信号(包括估计的提示词复杂度和模型置信度)映射为单查询采样预算。控制器为更简单或置信度更高的提示词分配更少的采样,为更困难或置信度更低的提示词分配更多采样,使推理时的计算资源可检查,而非固定或不透明。我们在公平对齐协议下(匹配解码设置和受控答案选择)进行评估,在问答和数学推理任务上与best-of-N、感知计算缩放及基于自置信度的基线进行比较。在不同模型和数据集上,自适应模糊控制优于多个标准基线,且在减少平均采样数的同时,仍接近与选择器匹配的全预算控制。这些发现表明,可解释自适应采样是实现更高效大语言模型测试时推理的实用方向。

英文摘要

Test-time scaling improves LLM reasoning by generating and aggregating multiple candidate answers, yet many pipelines use fixed per-query budgets that spend the same compute on easy and difficult prompts. These fixed budgets are also difficult to inspect because they do not explain why a given prompt receives a particular number of samples. We propose adaptive} test-time scaling with a lightweight fuzzy controller that maps interpretable signals, including estimated prompt complexity and model confidence, to a per-query sampling budget. The controller assigns fewer samples to easier or more confident prompts and more samples to harder or less certain prompts, making inference-time compute inspectable rather than fixed or opaque. We evaluate under a fair-alignment protocol with matched decoding settings and controlled answer selection, and compare against best-of-$N$, compute-aware scaling, and self-certainty-based baselines on question-answering and mathematical reasoning tasks. Across models and datasets, adaptive fuzzy control improves over several standard baselines and remains close to a selector-matched full-budget control while reducing the average number of samples. These findings suggest that interpretable adaptive sampling is a practical direction for more efficient test-time reasoning in large language models.

↑