arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34284cs.CL

过度个性化是一种决策失败:LLMs中生成引发的应用偏差

Over-Personalization Is a Decision Failure: Generation-Induced Apply Bias in LLMs

  • Chung-Ang University(中央大学)

机构由 AI 辅助整理,请以论文原文为准。

Haeun Jang, Yonghyun Jun, Hwanhee Lee

AI总结:

针对个性化LLM过度应用偏好问题,提出ABIDE方法,通过信号检测理论分析决策分数,发现生成目标引发应用偏差,并利用偏差标量校正以平衡满足与抑制。

AI中文摘要:

个性化大语言模型(LLMs)必须为每个存储的偏好决定当前上下文是否应应用或抑制该偏好,我们称之为其适用性。它们经常过度个性化,应用了上下文排除的偏好,然而现有基准仅对最终响应评分,无法判断这一失败发生在何处。我们将偏好处理分解为三个阶段并分别测量:(1) 知道偏好是否适用,(2) 决定显式的应用/抑制标签,以及(3) 生成与该标签一致的响应。使用线性探针,我们首先表明在生成过程中,该适用性信号仍可从隐藏状态中解码。通过使决策显式化,我们随后发现在大多数设置中,被忠实遵循的错误决策数量超过了在生成中丢失的正确决策。因此,我们将失败定位在决策环节,一旦模型也被要求回答,该环节即崩溃。为确定这反映的是敏感性丧失还是响应偏差,我们提出了ABIDE(应用偏差调查通过决策分数),它将信号检测理论应用于直接从logits读取的应用-抑制决策分数。ABIDE揭示了一种生成引发的应用偏差:仅仅陈述一个答案生成目标就会将决策分数移向应用,而敏感性基本保持不变,并且这种偏移在提示结构、跨偏好槽的级联以及提示措辞的控制下持续存在。最后,我们表明在解码时从决策分数中减去一个在保留分割上估计的单一偏差标量,可以在很大程度上保持满足的同时减少泄漏。

英文摘要:

Personalized LLMs must decide, for each stored preference, whether the current context calls for applying or suppressing it, which we call its applicability. They frequently over-personalize, applying preferences the context rules out, yet existing benchmarks score only the final response and cannot tell where this failure arises. We decompose preference handling into three stages and measure each separately: (1) knowing whether a preference applies, (2) deciding on an explicit Apply/Suppress label, and (3) generating a response consistent with that label. Using linear probes, we first show that this applicability signal remains decodable from hidden states during generation. By making the decision explicit, we then find that in most settings wrong decisions faithfully followed outnumber correct decisions lost in generation. We thus locate the failure in the decision, which breaks once the model is also asked to answer. To determine whether this reflects lost sensitivity or a response bias, we propose ABIDE (Apply-Bias Investigation via Decision-score), which adapts signal detection theory to Apply-vs-Suppress decision scores read directly from logits. ABIDE reveals a generation-induced Apply bias: merely stating an answer-generation objective shifts the decision score toward Apply while sensitivity is largely preserved, and the shift persists under controls for prompt structure, cascades across preference slots, and prompt wording. Finally, we show that subtracting a single bias scalar, estimated on a held-out split, from the decision score at decoding time reduces leakage while largely preserving fulfillment.

↑