BeliefScope:诊断大语言模型中证据驱动的修正与压力诱导的转变
BeliefScope: Diagnosing Evidence-Driven Revision and Pressure-Induced Shifts in Large Language Models
查看机构详情
- Beijing University of Posts and Telecommunications(北京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究提出BeliefScope框架,用于分离大语言模型中证据驱动的命题修正与压力诱导的响应转变,通过多维度测量与实验揭示其评估上下文依赖性及有效性边界。
中文摘要 AI 辅助
语言模型在接收到真正相关的证据后,或在接收到未添加相关事实的定向用户压力后,可能会对同一命题进行修正。因此,仅可观察到的响应转变无法揭示是哪一来源驱动了这一变化。我们引入BeliefScope,这是一个受控黑箱框架,用于围绕固定目标命题分离这两种影响来源。BeliefScope将证据与压力和特定因素的局部控制相结合,并通过概率报告、分类判断以及针对通道适配尺度的行动建议来测量响应变化。为确定这些可观察到的对比何时支持可靠归因,我们在受控合成条件下评估观察设计。已知真值恢复和针对性消融研究确定了可分离与证据和压力相关效应的区域,而半合成压力测试则绘制了随着观察过程变得更嘈杂和更异质,该可恢复性如何变化的情况。在一项包含36个模型系列的Qwen/Llama研究中,还有包含Gemma3-12B的针对性12个模型系列检查,所得轮廓显示出显著的评估上下文依赖性:广泛的模型级差异可在匹配控制、解码或响应界面下发生变化,而一些较窄的模型内模式保持稳定。指令干预进一步表明,减少的与目标对齐的压力依从性可反映稳定的抵抗或相反方向的转变。BeliefScope将这些测量总结为条件信念-响应轮廓,该轮廓将诊断效应与观察到它们的评估条件绑定在一起,同时附带该诊断的明确有效性边界。
英文摘要
A language model may revise the same proposition after receiving genuinely relevant evidence or after receiving directional user pressure that adds no relevant fact. The observable response shift alone therefore does not reveal which source drove the change. We introduce BeliefScope, a controlled black-box framework for separating these two sources of influence around a fixed target proposition. BeliefScope crosses Evidence and Pressure with factor-specific local controls and measures response changes through probability reports, categorical judgments, and action recommendations on channel-appropriate scales. To determine when these observable contrasts support reliable attribution, we evaluate the observation design under controlled synthetic conditions. Known-truth recovery and targeted ablations establish where Evidence- and Pressure-related effects can be separated, while semi-synthetic stress tests map how that recoverability changes as the observation process becomes noisier and more heterogeneous. Across a 36-family Qwen/Llama study, with targeted 12-family checks that also include Gemma3-12B, the resulting profiles show substantial evaluation-context dependence: broad model-level differences can change under matched controls, decoding, or response interfaces, while some narrower within-model patterns remain stable. Instruction interventions further show that reduced target-aligned Pressure following can reflect either stable resistance or movement in the opposite direction. BeliefScope summarizes these measurements as a conditional belief-response profile that keeps diagnostic effects tied to the evaluation conditions under which they are observed, together with explicit validity boundaries for that diagnosis.