arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14242cs.CL

通过概念链进行隐式推理引导

Implicit Reasoning Steering via Concept Chaining

Xiao Ye, Sanika Chavan, Yuxi Huang, Shahriar Kabir Nahin, Muhao Chen, Anshuman Chhabra, Ben Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

研究大语言模型推理脆弱性,提出概念链方法,通过生成连接段落继续预训练模型,使自然文本能引导模型预测,揭示推理脆弱性可被利用,为潜在偏差影响模型决策创造渠道。

中文摘要 AI 辅助

大语言模型看似能可靠推理,但重复采样时对许多问题会给出正确和错误答案,显示出最终决策形成存在潜在脆弱性。我们研究能否通过隐式推理引导利用这种脆弱性,即使用自然语言文本使模型偏向指定答案而无需明确指令等。我们的方法概念链生成短连接段落,通过一两个中间概念将问题实体与目标选项相连。接着在这些连接段落上继续预训练受害模型,评估其在原多选题上答案偏好是否改变。结果表明间接、自然的文本能系统引导模型预测,且比直接释义更难推断,说明推理脆弱性不仅是评估假象,还为潜在偏差通过普通文本放大并 covertly 重定向模型决策创造了实际渠道。

英文摘要

Large language models often appear to reason reliably, yet on many questions repeated sampling yields both correct and incorrect answers, revealing an underlying fragility in how final decisions are formed. We study whether this fragility can be exploited through implicit reasoning steering: using natural-language text to bias a model toward a designated answer without explicit instructions, triggers, or direct answer cues. Our approach, Concept Chaining, generates a short connection paragraph that links question entities to a target option through one or two intermediate concepts. We then continue pretraining a victim model on these connection paragraphs and evaluate whether its answer preference shifts on the original multiple-choice questions. Our results show that indirect, natural-looking text can systematically steer model predictions while remaining substantially less inferable than direct paraphrases, which shows that reasoning brittleness is not merely an evaluation artifact: it creates a practical channel through which latent biases can be amplified by ordinary-looking text to covertly redirect model decisions.

发表机构

  • School of Computing and Augmented Intelligence, Arizona State University(亚利桑那州立大学计算与增强智能学院)
  • Department of Computer Science, University of California, Davis(加利福尼亚大学戴维斯分校计算机科学系)

机构由 AI 辅助整理,请以论文原文为准。

↑