arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可解码性准则预测大型语言模型中隐藏状态选择何时优于多数投票

A decodability criterion predicts when hidden-state selection beats majority voting in large language models

Zhixiang wang, Ziliang Hong, Ulas Bagci

arXiv 2608.17124首次发表:更新:

发表机构

Northwestern University(西北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出可解码性准则及CASE方法,可预测大型语言模型中隐藏状态选择何时优于多数投票,CASE在通用和医疗LLM的中、高难度问题上均有显著提升,且预测可跨领域迁移。

AI 中文摘要

将大型语言模型(LLM)针对某个问题采样得到的多个答案合并为一个决策是一种测试时信息融合问题,通常采用多数投票法解决。在困难问题上,投票法不可靠,因为采样得到的答案存在相关误差,可能导致错误答案胜出,且增加采样数量会使决策更差。从模型隐藏状态中读取正确性信号来选择候选答案是一种有前景的替代方案,但其准确率因模型和任务而异,且没有指标可表明何时可信任该方案。本文提出CASE(Correctness-Axis SElection,正确性轴选择),一种动态选择组合器,它在答案标记的隐藏状态上训练线性门,并选择得分最高的候选答案。其主要贡献是可解码性,这是一种无泄漏的指标,用于衡量门对问题正确候选答案的排序优于错误候选答案的程度,可预测隐藏状态选择是否会优于投票法。传统探针看似准确,实则仅因问题身份泄漏导致,在按问题分组的评估下该泄漏会消失。在保留的数据集上,可解码性以皮尔逊相关系数r=0.75和接近AUC=0.60的决策阈值,预测选择法相对于投票法的准确率提升。在通用和医疗LLM上,CASE在中等难度问题上较投票法提升最高19个百分点,在困难问题上提升16.8个百分点。可解码性取决于模型必须回忆的对齐知识,而非模型规模,其预测可在3.8个百分点内迁移至未见过的科学领域。因此,它提供了一种实用准则,可针对给定模型和任务提前测量,用于在学习型选择法和多数投票法之间做出选择。

英文摘要

Combining the answers a large language model (LLM) samples for a question into one decision is a test-time information fusion problem, usually solved by majority voting. Voting is unreliable on difficult questions, where the sampled answers share correlated errors, so the wrong answer can win and drawing more samples makes the decision worse. Selecting a candidate by reading a correctness signal from the model's hidden states is a promising alternative, but its accuracy varies across models and tasks, and no measure indicates when it can be trusted. In this paper, we propose CASE (Correctness-Axis SElection), a dynamic selection combiner that trains a linear gate on the answer-token hidden state and selects the highest-scoring candidate. Its main contribution is decodability, a leakage-free measure of how well the gate ranks a question's correct candidates above its incorrect ones, which predicts whether hidden-state selection will outperform voting. A conventional probe appears accurate only because of question-identity leakage, which vanishes under question-grouped evaluation. On held-out data, decodability predicts the accuracy gain of selection over voting with a Pearson correlation r=0.75 and a decision threshold near AUC=0.60. Across general and medical LLMs, CASE improves over voting by up to 19 points on medium-difficulty questions and 16.8 points on hard questions. Decodability depends on the aligned knowledge a model must recall, not on its scale, and its prediction transfers to an unseen scientific domain within 3.8 points. It thus provides a practical criterion, measurable in advance for a given model and task, for choosing between learned selection and majority voting.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑