arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

神经解码中上下文先验下的置信度排序反转

Confidence-Ordering Reversal under Contextual Priors in Neural Decoding

Xinyu Zhang, Sichao Liu

arXiv 2610.08229首次发表:更新:

发表机构

KTH; Karolinska Institutet(瑞典皇家理工学院; 卡罗林斯卡学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示神经解码中上下文先验导致置信度排序反转,提出分离读取局部与先验分数以改善选择性解码,使回答窗口从56.7%提升至74.5%。

AI 中文摘要

上下文先验通过重塑候选分数来改善神经到语言的解码。然而,置信度是从同样的重塑分数中读取的,因此先验留下的错误可能变得更加自信,而准确性没有变化来揭示这一点。我们研究了在MEG-MASC和MOUS数据集上,使用局部解码分数、通过加性浅融合结合的上下文先验以及融合后的前二名间隔作为置信度,先验如何塑造语音检索中的置信度。在最初错误的预测中,我们发现了一种置信度排序反转:当正确候选在局部排名中起始位置接近顶部时,更大的间隔使修复更可能发生;但当其起始位置较低时,更大的间隔使修复更不可能。在MEG-MASC上,合并的正确性AUROC为0.87,但区分修复与残余错误的AUROC从初始排名2-3的0.70下降到排名21-50的0.39。起始于排名20之后、位于反转区域内的错误,占所有融合后错误的46.6%。我们提出了一个分数层面的解释:修复必须首先缩小正确候选的初始差距,限制其最终间隔,而残余错误可以在两个错误候选之间建立大的间隔。一个仅改变融合权重的因果干预,按预测将反转移动到更深的排名。在词级LM先验下,它在准确性增益达到峰值后继续移动,因此为准确性选择的权重并不能稳定置信度。分别读取局部和先验分数改善了选择性解码:解码器在74.5%的窗口上回答,而不是56.7%,而92%的输出集仍然包含正确候选。上下文融合后的置信度应保留每个预测背后的局部和上下文证据,而不仅仅是融合后的分数。项目网站:此https链接 代码:此https链接

英文摘要

Contextual priors improve neural-to-language decoding by reshaping candidate scores. However, confidence is read from the same reshaped scores, so the errors a prior leaves behind can become more confident with no change in accuracy to reveal it. We study how a prior shapes confidence in speech retrieval on MEG-MASC and MOUS using local decoding scores, a contextual prior combined by additive shallow fusion, and the fused top-two margin as confidence. Among initially incorrect predictions, we find a confidence-ordering reversal: a larger margin makes a repair more likely when the correct candidate starts near the top of the local ranking, but less likely when it starts lower. On MEG-MASC, pooled correctness AUROC is 0.87, yet AUROC separating repairs from residual errors falls from 0.70 at initial ranks 2-3 to 0.39 at ranks 21-50. Errors starting beyond rank 20, inside the reversed region, make up 46.6% of all post-fusion errors. We propose a score-level account: a repair must first close the correct candidate's initial deficit, limiting its final margin, whereas a residual error can build a large margin between two incorrect candidates. A causal intervention that changes only the fusion weight moves the reversal to deeper ranks as predicted. Under a word-level LM prior, it keeps moving after accuracy gain peaks, so a weight chosen for accuracy does not settle confidence. Reading local and prior scores separately improves selective decoding: the decoder answers on 74.5% of windows instead of 56.7%, while 92% of output sets still contain the correct candidate. Confidence after contextual fusion should retain the local and contextual evidence behind each prediction, not just the fused scores. Project website: https://confidencereversal.github.io/; Code: https://github.com/AmadeusFake/NeuDecodingConfReversal

Comments28 pages, 4 figures, 18 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑