发表机构
Harbin Institute of Technology, Shenzhen; Huawei Technologies Co., Ltd.; Beijing Academy of Artificial Intelligence(哈尔滨工业大学(深圳); 华为技术有限公司; 北京智源人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对推理时LLMs密集对齐的低置信度干预问题,提出TUSA方法,通过不确定性感知仲裁器仅在满足条件时授权干预,跳过约50%对齐步骤,同时提升安全与通用偏好。
AI 中文摘要
推理时对齐的一个突出范式是使用轻量监督器引导大型语言模型(LLMs)。通过实证分析,我们发现该范式存在结构不匹配:弱监督器在绝大多数token上呈现普遍的高熵,而主流的密集干预方法要求在每一步解码时都进行监督,这导致频繁的低置信度干预,可能破坏基础模型的有效推理并产生大量实用成本。为解决该问题,我们提出TUSA(基于信任的不确定性稀疏对齐)。TUSA摒弃持续监督,将对齐重新定义为动态仲裁过程,引入不确定性感知仲裁器,仅在满足两个条件时授权干预:监督器具有置信度,且该token在语义上具有显著性。该机制有效过滤了不确定性驱动的噪声和冗余监督。在多个模型和基准上的大量实验表明,TUSA始终能同时提升安全对齐和通用有用性;与密集基线相比,它跳过约50%的对齐步骤,不仅将安全偏好提升了最高15.6%,还将通用偏好率提升了最高12.0%,证明选择性的高精度对齐优于持续监督。
英文摘要
A prominent paradigm in inference-time alignment employs lightweight supervisors to steer Large Language Models (LLMs). Through empirical analysis, we identify a structural mismatch in this paradigm: weak supervisors exhibit pervasive high entropy across the vast majority of tokens, yet prevailing dense intervention approaches mandate supervision at every decoding step. This leads to frequent low-confidence interventions that can disrupt valid base-model reasoning and incur substantial utility costs. To resolve this, we propose TUSA (Trust-based Uncertainty Sparse Alignment). Moving away from continuous oversight, TUSA reframes alignment as a dynamic arbitration process, introducing an uncertainty-aware arbiter that authorizes intervention only when two conditions are met: the supervisor is confident and the token is semantically salient. This mechanism effectively filters out uncertainty-driven noise and redundant supervision. Extensive experiments across multiple models and benchmarks show that TUSA consistently improves both safety alignment and general helpfulness. By bypassing approximately 50% of alignment steps, it not only enhances safety preference by up to 15.6%, but also boosts general preference rates by up to 12.0% compared to the dense baseline, demonstrating that selective, high-precision alignment can outperform continuous supervision.
CommentsAccepted to Findings of EMNLP 2026