arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05166cs.CLcs.CYcs.HCcs.LG

大语言模型中的条件认知偏差:有偏差的用户对话轮次如何调控上下文推理

Conditional Cognitive Biases in LLMs: How Biased User Turns Modulate In-Context Reasoning

Sachini Weerasekara, Sagar Kamarthi, Jacqueline Isaacs

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过三条件实验框架与含24300条提示的基准数据集,评估8个前沿LLMs在多轮交互中受有偏差用户轮次的影响,发现多数模型偏差表现变化,明确相关动态并发布研究资源。

中文摘要 AI 辅助

我们在现实多轮交互场景下,对最先进的指令微调大语言模型(LLMs)的认知偏差表现进行了评估。本研究引入了一种新颖的三条件实验框架,该框架可将接触有偏差的用户对话轮次的影响与该轮次语义内容的影响分离开来,同时还构建了一个包含24300条经评审团验证的用户提示的基准数据集,这些提示覆盖了9×9目标人类偏差交互矩阵的全部81个单元。在8个前沿LLMs中,我们发现,与零样本基线相比,有偏差的对话上下文会系统性地增加8个模型中6个模型的偏差表现。我们确定了导致该效应的两种相互竞争的行为动态:接触有偏差的推理通常会放大下游偏差倾向,而明确表述的偏差线索往往会触发与对齐相关的抑制行为,从而减少显性偏差表现。我们发布了该框架、代码库和数据集,以支持未来关于LLMs中条件认知偏差和行为适应的研究。

英文摘要

We present an evaluation of cognitive bias expression in state-of-the-art instruction-tuned LLMs under realistic multi-turn interaction settings. Our work introduces a novel three-condition experimental framework that disentangles the effect of exposure to a biased user turn from the effect of the turn's semantic content, alongside a benchmark of 24,300 jury-validated user prompts spanning all 81 cells of a 9x9 target-human bias interaction matrix. Across eight frontier LLMs, we find that biased conversational context systematically increases bias expression relative to zero-shot baselines in 6 of 8 models. We identify two competing behavioral dynamics underlying this effect: conversational exposure to biased reasoning generally amplifies downstream bias tendencies, while explicitly stated bias cues often trigger alignment-related suppression behaviors that reduce overt bias expression. We release our framework, codebase, and dataset to support future research on context-conditioned cognitive biases and behavioral adaptation in LLMs.

发表机构

  • Northeastern University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

↑