arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OptimismBench:语言模型判断中的偏差预测与对齐效应

OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment

Seonglae Cho, Adriano Koshiyama

arXiv 2607.26981首次发表:更新:

发表机构

Holistic AI; University College London(整体人工智能公司; 伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出OptimismBench基准,检测到多数LLM存在乐观方向偏差,且对齐会使模型概率产生偏差,下游流程会默认继承该偏差。

AI 中文摘要

大型语言模型(LLM)越来越多地被用作决策辅助工具,其概率判断会影响下游选择。这些判断是否存在系统性的方向偏差一直难以检测:校准指标汇总的是无符号误差,而自然不确定性没有真实概率。当LLM对一家初创企业的成功概率给出70%、失败概率给出15%时,缺失的15个百分点暴露了任何汇总分数都未标记的偏差。我们推出OptimismBench,该基准通过反向配对检测方向偏差:每个场景会同时引出P(成功)和P(失败),两种框架之间的不对称性可在无需真实概率的情况下得出有符号偏差分数。在来自8个提供商的16个模型中,14个表现出乐观偏差;悲观偏差仅出现在Anthropic的前沿层级模型中。四个模型家族的11组基础模型与对话模型配对显示,训练后的模型会确定偏差的符号,且不同家族会出现相反方向的偏差变化。该模式在提示、温度、视角和自我去偏的 ablation 实验中依然存在。17个模型、6种语言的对比进一步显示,模型身份的影响远大于语言的影响,模型间方差是语言间方差的4.7倍。我们发布了涵盖10种语言的3870个条目,用于每个模型的方向偏差审计。当对齐使模型更具帮助性时,也会使其概率产生偏差;下游流程会默认继承这种偏差。

英文摘要

Large language models are increasingly used as decision aids whose probability judgments shape downstream choices. Whether those judgments carry a systematic directional tilt has been hard to detect: calibration metrics aggregate unsigned errors, and naturalistic uncertainty offers no ground-truth probability. When an LLM rates a startup's success at 70% but its failure at 15%, the missing 15 points expose a distortion no aggregate score flags. We introduce OptimismBench, which detects directional bias with inverted pairs: each scenario elicits both P(success) and P(failure), and asymmetry between the two framings yields a signed bias score without ground truth. Across 16 models from 8 providers, fourteen are optimistic; pessimism appears only in Anthropic's frontier tier. Eleven matched base-versus-chat pairs across four families show post-training sets the sign of the bias, with opposite shifts in different families. The pattern survives prompt, temperature, perspective, and self-debiasing ablations. A seventeen-model six-language comparison further shows model identity dominates language, with inter-model variance at 4.7x inter-language variance. We release 3,870 items across 10 languages for per-model directional-bias auditing. When alignment makes a model more helpful, it also tilts its probabilities; downstream pipelines inherit the tilt by default.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑