arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

仅在确信时信任你的引导:推理时的不确定性感知稀疏对齐

Trust Your Guide Only When Certain: Uncertainty-Aware Sparse Alignment at Inference Time

Zeen Zhu, Zhuo Li, Weiyang Guo, Liye Zhao, Haibing Di, Yequan Wang, Jing Li

arXiv 2609.00624首次发表:更新:

发表机构

Harbin Institute of Technology, Shenzhen; Huawei Technologies Co., Ltd.; Beijing Academy of Artificial Intelligence(哈尔滨工业大学(深圳); 华为技术有限公司; 北京智源人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对推理时LLMs密集对齐的低置信度干预问题,提出TUSA方法,通过不确定性感知仲裁器仅在满足条件时授权干预,跳过约50%对齐步骤,同时提升安全与通用偏好。

AI 中文摘要

推理时对齐的一个突出范式是使用轻量监督器引导大型语言模型(LLMs)。通过实证分析,我们发现该范式存在结构不匹配:弱监督器在绝大多数token上呈现普遍的高熵,而主流的密集干预方法要求在每一步解码时都进行监督,这导致频繁的低置信度干预,可能破坏基础模型的有效推理并产生大量实用成本。为解决该问题,我们提出TUSA(基于信任的不确定性稀疏对齐)。TUSA摒弃持续监督,将对齐重新定义为动态仲裁过程,引入不确定性感知仲裁器,仅在满足两个条件时授权干预:监督器具有置信度,且该token在语义上具有显著性。该机制有效过滤了不确定性驱动的噪声和冗余监督。在多个模型和基准上的大量实验表明,TUSA始终能同时提升安全对齐和通用有用性;与密集基线相比,它跳过约50%的对齐步骤,不仅将安全偏好提升了最高15.6%,还将通用偏好率提升了最高12.0%,证明选择性的高精度对齐优于持续监督。

英文摘要

A prominent paradigm in inference-time alignment employs lightweight supervisors to steer Large Language Models (LLMs). Through empirical analysis, we identify a structural mismatch in this paradigm: weak supervisors exhibit pervasive high entropy across the vast majority of tokens, yet prevailing dense intervention approaches mandate supervision at every decoding step. This leads to frequent low-confidence interventions that can disrupt valid base-model reasoning and incur substantial utility costs. To resolve this, we propose TUSA (Trust-based Uncertainty Sparse Alignment). Moving away from continuous oversight, TUSA reframes alignment as a dynamic arbitration process, introducing an uncertainty-aware arbiter that authorizes intervention only when two conditions are met: the supervisor is confident and the token is semantically salient. This mechanism effectively filters out uncertainty-driven noise and redundant supervision. Extensive experiments across multiple models and benchmarks show that TUSA consistently improves both safety alignment and general helpfulness. By bypassing approximately 50% of alignment steps, it not only enhances safety preference by up to 15.6%, but also boosts general preference rates by up to 12.0% compared to the dense baseline, demonstrating that selective, high-precision alignment can outperform continuous supervision.

CommentsAccepted to Findings of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑