arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01081cs.CLcs.AI

StateSwap:探查多项选择题中支持-消除式的隐藏状态

StateSwap: Probing Support-Elimination Hidden States in Multiple-Choice Questions

  • Vrije Universiteit Amsterdam(阿姆斯特丹自由大学)
  • Wuhan University of Science and Technology(武汉科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Chao Gao, Haijiang Liu, Qiyuan Li, Caicai Guo, Frank van Harmelen, Jinguang Gu

AI总结:

本研究针对大型语言模型在多项选择题不同框架下的答案不一致问题,提出双框架协议,通过交换[STATE]标记的残差流激活值,验证其与模型行为相关,还发现均值差引导方向的层内响应更具边界性。

AI中文摘要:

大型语言模型在面对支持导向和消除导向两种不同框架下提出的同一道多项选择题时,常常给出不一致的答案。我们探究这种差异是否源于两种框架诱导出的不同内部表征。我们引入了一种双框架协议,该协议使用仅在支持导向或消除导向框架上有细微差异的提示,同时保持评估目标固定不变。为探查内部计算,我们添加了一个未经过训练的特殊标记[STATE],并将其残差流激活值作为干预接口。在两种模型中,两种框架都诱导出了可分离的[STATE]激活值,且这些激活值集中在中间层。在配对提示之间交换这些激活值会系统性地改变预测结果,并提升跨框架一致性,这提供了基于干预的证据,表明这些激活值具有行为相关性。除了实例级替换外,在评估协议下,从双框架对比中得出的均值差引导方向,比匹配的对比激活添加方向表现出更具边界性的层内响应。

英文摘要:

Large language models often answer the same multiple-choice question inconsistently when it is posed under support-oriented and elimination-oriented framings. We investigate whether these discrepancies arise from different internal representations induced by the two framings. We introduce a dual-framing protocol with minimally varied prompts that use either support- or elimination-oriented framing while keeping the evaluation target fixed. To probe the internal computation, we append an untrained special token, [STATE], and treat its residual-stream activation as an intervention interface. Across both models, the two framings induce separable [STATE] activations concentrated in intermediate layers. Swapping these activations between paired prompts systematically changes predictions and improves cross-framing agreement, providing intervention-based evidence that the activations are behaviorally relevant. Beyond instance-level substitution, mean-difference steering directions derived from the dual-framing contrast exhibit more bounded layer-wise responses than matched contrastive activation addition directions under the evaluated protocol.

补充信息

↑