大语言模型风险决策中自我报告与行为一致性的表征控制
Representational Control over Self-Report & Behavior Coherence in LLM Risk-Taking
浏览论文内容
中文总结 AI 辅助
本研究通过激活引导干预,发现大语言模型自我报告与风险行为由近乎正交的内部表征通道控制,组合干预可同时影响两者,翻转符号则使其对立,揭示了二者关系的表征基础。
中文摘要 AI 辅助
自我报告是探测大语言模型倾向的一种低成本手段,但近期研究发现模型所报告的内容与其实际行为之间仅存在选择性一致。以往研究通过提示黑盒大语言模型来确立这些模式,因此尚不清楚这种差异究竟是提示产生的伪影,还是底层构念在内部表征方式上的事实。我们以风险承担——智能体决策中一个具有重要影响的维度——为研究对象,利用激活引导在同一内部干预下测量自我报告与行为。我们调查了九种引导向量提取方法,涵盖任务特定指令、模型自身的任务行为以及两种粒度下的倾向性描述,并在四个开放权重大语言模型上,通过两项行为任务和两项心理测量工具进行评估。我们发现:(1)共享的内部干预并不能确保共享的响应性:由特质描述构建的方向能改变自我报告,但行为仍处于随机水平;由模型自身任务选择构建的方向则相反;只有任务特定指令能同时影响两者,但效果较弱。(2)诊断分析消除了这一例外:去除表面混杂因素后,指令的行为效果仅保留32%。两个通道在其他情况下由近乎正交的方向引导,每个通道均可通过多种独立构建方式达到。(3)将一条行为移动方向和一条自我报告移动方向组合的干预能同时移动两者;翻转其中一个符号则使两者对立,在所有模型中,69-89%的翻转组合中报告的风险与执行的风险指向相反方向。这些发现将自我报告与行为的关系从黑盒观察提升为可检查和可控的表征层面关系,促使在行为评估之外进行表征检查。
英文摘要
Self-report is an appealing low-cost probe of an LLM's dispositions, but recent work finds only selective agreement between what models report and how they behave. Prior accounts establish these patterns by prompting black-box LLMs, leaving open whether the gap is a prompting artefact or a fact about how the underlying constructs are represented internally. We investigate risk-taking, a consequential dimension of agentic decision-making, using activation steering to measure self-report and behavior under the same internal intervention. We survey nine steering-vector extraction methods spanning task-specific directives, the model's own task behavior, and dispositional descriptions at two granularities, evaluated on two behavioral tasks and two psychometric instruments across four open-weight LLMs. We find that (1) a shared internal intervention does not ensure shared responsiveness: directions built from trait descriptions move self-report but leave behavior at chance, directions built from the model's own task choices do the reverse, and only task-specific directives reach both, weakly. (2) Diagnosis dissolves that exception: removing surface confounders leaves the directives only 32% of their behavioral effect. The two channels are otherwise steered by near-orthogonal directions, each channel reachable by several independent constructions. (3) An intervention composing one behavior-moving and one self-report-moving direction moves both together; flipping one sign sets them in opposition, with reported and enacted risk pointing in opposite directions on 69-89% of flipped compositions in all models. These findings move the self-report-behavior relationship from a black-box observation to a representational one that can be inspected and controlled, motivating representational checks alongside behavioral evaluation.
发表机构
- Caltech(加州理工学院)
- Carnegie Mellon University(卡内基梅隆大学)
- UIUC(伊利诺伊大学厄巴纳-香槟分校)
- University of Cambridge(剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。