DUD:用于大语言模型可靠不确定性量化的解耦更新动力学
DUD: Decoupled Update Dynamics for Reliable Uncertainty Quantification in Large Language Models
浏览论文内容
中文总结 AI 辅助
本研究针对大语言模型不确定性量化的传统方法缺陷,提出DUD框架通过因果干预解耦FFN与注意力的更新贡献,实验表明其在不确定性估计、校准及跨数据集泛化上均优于现有基线。
中文摘要 AI 辅助
准确的不确定性量化(UQ)对大语言模型(LLM)的可靠部署至关重要,但传统基于概率的指标往往无法捕捉模型的真实认知状态。尽管近期的机制方法利用了隐藏状态动力学,却通常会聚合残差流更新,混淆了参数记忆(前馈网络,FFN)与上下文处理(注意力机制)的不同作用。我们认为这种聚合会掩盖细粒度的机制冲突,比如记忆-上下文失配,而这正是不确定性的核心指标。为解决该问题,我们提出DUD(Decoupled Update Dynamics,解耦更新动力学)框架,该框架通过噪声诱导的因果干预明确解耦FFN与注意力机制的贡献。通过量化每个模块的独立恢复能力,我们构建了双路动态轮廓以捕捉模型的内部脆弱性。大量实验表明,DUD在不确定性估计与校准方面均显著优于当前最优基线,同时具备更出色的跨数据集泛化能力,验证了解耦动力学可作为模型忠实度的可靠替代指标。
英文摘要
Accurate Uncertainty Quantification (UQ) is critical for reliable deployment of Large Language Models (LLMs), yet traditional probability-based metrics often fail to capture the model's true epistemic state. While recent mechanistic approaches leverage hidden state dynamics, they typically aggregate residual stream updates, conflating the distinct roles of parametric memory (Feed-Forward Networks) and contextual processing (Attention). We argue that this aggregation obscures fine-grained mechanistic conflicts, such as memory-context misalignment, that are fundamental indicators of uncertainty. To address this, we introduce \textbf{D}ecoupled \textbf{U}pdate \textbf{D}ynamics \textbf{(DUD)}, a framework that explicitly decouples FFN and Attention contributions via noise-induced causal interventions. By quantifying the independent restoration capabilities of each module, we construct a dual-stream dynamic profile that captures the model's internal fragility. Extensive experiments demonstrate that DUD significantly outperforms state-of-the-art baselines in both uncertainty estimation and calibration, while exhibiting superior cross-dataset generalization, validating decoupled dynamics as a robust proxy for model faithfulness.