arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

跨模型价值比较中的两个混淆因素:响应确定性和访问约束

When Counterbalancing Hides the Bias: Access-Conditioned Position Lock in Forced-Choice LLM Evaluation

Hong-In Won, Jinseok Jang, Hyoseop Kim

arXiv 2607.10202首次发表:更新:

发表机构

KITECH; National Research Council of Science and Technology (NST)(韩国产业技术评价院; 韩国国家科学技术研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究跨模型价值比较中的两个混淆因素,提出无规则价值困境等方法分离并校正响应确定性,还发现访问约束影响模型价值概况,贡献了确定性校正分解,识别出部署约束是独特价值混淆因素。

AI 中文摘要

跨模型比较将价值倾向的差异视为语言模型具有个性化价值的证据。在单次抽取测量下,这混淆了两个量:集中趋势的差异(真正的价值差异)和响应确定性的差异(模型对强制选择的坚定程度)。我们引入了一种分离协议——无规则价值困境,通过平衡、重复的强制选择测量和确定性指数——以及一种确定性校正分解,将明显的跨模型距离分解为方向翻转分量(真正的分歧)和同侧更极端分量(我们称为确定性)。在九个模型中,确定性差异很大(参与模型中为0.66 - 0.95);它是每个模型的特征还是跟踪提供者和规模,我们的方法使其可测量,但样本未明确。校正确定性会缩小明显的个性化,同时一些跨家族分歧在严格测试中幸存。然后我们分离出第二个混淆因素:服务于每个模型的访问约束。通过原始提供者API重新收集相同模型,我们发现部署客户端会显著且特定于客户端地改变模型的价值概况:一个订阅CLI使概况移动0.31,翻转18个项目中的4个,并增加旗舰产品的明显柔软度(通过CLI为0.34,通过原始API为0.66),而另一个提供者的客户端则无此问题,这混淆了提供者家族和访问客户端。约束是一个价值塑造层:一个拒绝十分之一强制选择的基础模型通过代理系统提示变得合规,在白盒控制中因果确定。因此,按单次抽取价值距离对模型进行审计排名时,排名的是一个因确定性而膨胀的量,还会因使用的客户端而进一步混淆。我们贡献了这种分解,并将部署约束识别为一个独特的价值混淆因素。

英文摘要

Forced-choice probes with counterbalanced orientations are a standard tool for measuring language-model "value dispositions," and a concentration/extremity index over repeated draws is read as how sharply a model commits. We show this estimator is not identifiable at its low end: counterbalancing, meant to remove position bias, instead maps a position-lock (a model returning the same answer letter regardless of content) onto the same near-0.5 signature as genuine neutrality, so "softness" and "non-engagement" cannot be distinguished from a content-independent letter-bias. Across nine models the fraction of position-locked items tracks the index almost perfectly (r=-0.986) - a structural consequence, not a finding: the informative quantity is the residual from that bound, the commitment a model shows on the items it does engage. The three models the index reads as soft are the most position-locked, and the lock resolves under access configurations that permit reasoning (as-deployed CLI -> raw API -> reasoning-enabled), moving the index 0.06->0.63 and 0.34->0.66->0.64 while lock collapses (Opus: 0.61->0.39->0.22); this establishes the low readings as an identifiability failure, not a disposition. The reasoning-enabled reading is not a "true" value either; the point is that the index alone cannot identify commitment at its low end. Access configuration (deployment client, reasoning on/off) is one generator of this failure, shown on two access paths (an Anthropic subscription CLI and a DeepSeek client), and in a frozen benchmark it is confounded with provider. We contribute a position-lock diagnostic that must accompany any concentration reading, and show that a concentration-blind audit risks reporting neutrality where a reasoning-permitting condition yields concentrated choices. The direction-flip component, largely robust to the artifact, still identifies genuine cross-model disagreement.

Comments16 pages, 3 figures, 7 tables. Substantially revised: reframed around an identifiability result for the counterbalanced concentration index, with access configuration presented as one generator of the failure rather than a separate confound. Code and data are included as ancillary files

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑