认知立场灵活性探测:测量大语言模型中提示条件下的语域转换
Epistemic Stance Flexibility Probing: Measuring Prompt-Conditioned Register Shift in Large Language Models
浏览论文内容
中文总结 AI 辅助
研究大语言模型在不同归因提示下的认知立场灵活性,引入ESFP基准,涵盖多类别多模板项目,从四个维度评估模型响应,发现认知灵活性与模型能力正交,立场内容密度信号最强。
中文摘要 AI 辅助
语言模型可能会被问到专家对有争议的主张的看法,或者它自己对该主张的看法。一个值得信赖的对话代理应该区分这两个请求,并以不同的认知语域做出回应。现有的准确性、指令跟随或安全性基准并未直接评估这种转换是否发生以及是否连贯。我们引入了ESFP,一个行为基准,将外部归因提示和自我归因提示之间的对比作为基本测量单位。ESFP由104个精心控制的项目组成,跨越六个认知类别和五个措辞模板,并从四个互补维度评估模型响应:词汇自我归因、对角色框架的表示级响应、由大语言模型评判小组评估的句子级立场内容密度以及跨条件立场一致性。评估来自五个供应商的八个前沿模型后发现,认知灵活性在很大程度上与一般模型能力正交。立场内容密度提供了最强的信号。我们提供了项目级自举置信区间、权重敏感性分析以及对综合分数解释限制的明确讨论。ESFP测量模型在变化的归因条件下调整其认知立场的倾向,而不是一般的能力测量。
英文摘要
A language model may be asked either what experts believe about a contested claim or what it believes about the claim itself. A trustworthy conversational agent should distinguish these two requests and respond in different epistemic registers: neutral attribution in the first case and stance expression in the second. Whether such a shift occurs-and whether it occurs coherently-is not directly assessed by existing benchmarks for accuracy, instruction following, or safety. We introduce ESFP, a behavioral benchmark that treats the contrast between externally attributed and self-attributed prompts as the fundamental unit of measurement. ESFP consists of 104 carefully controlled items spanning six epistemic categories and five phrasing templates, and evaluates model responses along four complementary dimensions: lexical self-attribution, representation-level responsiveness to role framing, sentence-level stance content density assessed by an LLM judge panel, and cross-condition stance consistency. Evaluating eight frontier models from five vendors, we find that epistemic flexibility is largely orthogonal to general model capability: a 27B open-weight model matches the strongest proprietary systems, the flagship model of one family underperforms its lightweight counterpart, and reasoning-optimized models do not consistently exhibit higher flexibility. Stance content density provides the strongest signal, while surface-level lexical markers such as 'I think' can change substantially without corresponding changes in expressed stance. We provide item-level bootstrap confidence intervals, weight-sensitivity analyses, and an explicit discussion of the interpretation limits of the composite score. ESFP measures a model's propensity to adapt its epistemic stance under changing attribution conditions, rather than a general competence measure.