arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07117cs.CL

去偏的幻觉:人物角色引导在LLMs中重新分配而非减少偏见

The Illusion of Debiasing: Persona Steering Redistributes Rather Than Reduces Bias in LLMs

Ziyue Feng, Hongbo Fang, James A. Evans

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过人物角色条件探针发现,提示干预仅调节输出通道而非内部结构,可重新分配但无法减少模型偏见,且存在结构性触及极限。

中文摘要 AI 辅助

基于提示的干预措施(如系统提示、人物角色、角色指令)能够可靠地重塑语言模型的输出内容,但其作用层次尚不明确。这些干预措施是重新配置了内部结构,还是仅仅调节了输出通道?我们以人物角色条件作为受控探针,沿着从自我报告、开放式生成到词级参数关联的深度轴测量其影响,涉及三个经过指令微调的模型。我们发现了一种分级分离现象。人物角色是可读的但非结构性的:模型遵循单一特质的指令,却无法复现人类特质间的协方差。这种分离随深度加深而加剧:人物角色维持或放大封闭式问答中的偏见,改变绝对语气但群体间差异保持不变,并且几乎不扰动已饱和的联想基线。因此,基于提示的引导作用于输出通道,其结构性触及范围存在限制,而这种限制可能被表面可操作性所掩盖。

英文摘要

Prompt-based interventions: system prompts, personas, role instructions, reliably reshape what a language model says, but it is unclear which layer they reach. Do they reconfigure internal structure, or only modulate the output channel? We use persona conditioning as a controlled probe, measuring its effects along a depth axis from self-report, through open-ended generation, to word-level parametric association, across three instruction-tuned models. We find a graded dissociation. Personas are legible but not structural: models follow single-trait instructions yet fail to reproduce human inter-trait covariance. The dissociation deepens with depth: personas hold or amplify closed-form QA bias, shift absolute tone while leaving between-group disparity unchanged, and barely perturb an already saturated associative baseline. Prompt-based steering thus operates in the output channel and has a structural reach limit that surface manipulability can mask.

发表机构

  • University of Chicago(芝加哥大学)

机构由 AI 辅助整理,请以论文原文为准。

↑