arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

论点与信头:AI评估中的来源-立场一致性

The Argument and the Letterhead: Source-Position Coherence in AI Evaluation

Michele Loi

arXiv 2609.35286首次发表:更新:

AI 中文总结

本研究通过受控比较实验,发现AI评估者能区分来源与论点质量,支持来源-立场一致性解释,并报告了相关交互作用及事后p值。

AI 中文摘要

一个论点如果出自特定发言人之口,可能会令人惊讶,但这并不妨碍它本身是一个好论点。AI评估者能否区分这两种判断?两项预注册的描述性研究和一项后续的Jev补充研究收集了2,976份针对六篇固定文本的有效评估,这些文本涉及美国AI政策、德国债务刹车和瑞士核能。每篇文本在多个来源归属下呈现。关键比较在于,当论点改变时,两个来源之间的差距是否发生变化。例如,在Sol上,一项国家安全论点在CODEPINK下获得平均评分0.359,在College Republicans下获得0.639;一项民权论点分别获得0.742和0.721。对某一来源的恒定偏好无法解释这种模式。相关交互作用出现在不同主题和最新模型配置中,包括启用推理的配置,而若干比较产生了较小效应。后续的欧洲Jev补充研究产生了五项低于采用的绝对参考值0.05的交互作用;其独特的评分标准和中断的收集过程限制了对聊天系统的比较。一些书面评估明确提及来源与其被归属立场之间的不匹配。综合来看,数值和文字证据支持来源-立场一致性作为合理解释,同时存在涉及可信度、真实性和任务解读的竞争性解释。本文通过受控比较发展这一推断,在附录中报告了条件性事后p值,并记录了AI主导研究背后的人类决策和委派检查。

英文摘要

An argument can be surprising coming from a particular speaker without being a bad argument. Do AI evaluators keep these judgments apart? Two preregistered descriptive studies and a later Jev supplement collected 2,976 usable evaluations of six fixed texts about US AI policy, Germany's debt brake and Swiss nuclear energy. Each text was presented under several source attributions. The key comparison asks whether the gap between two sources changes when the argument changes. On Sol, for example, a national-security argument received mean ratings of 0.359 under CODEPINK and 0.639 under College Republicans; a civil-rights argument received 0.742 and 0.721. A constant preference for one source cannot explain that pattern. Related interactions appeared across topics and recent model configurations, including those with reasoning enabled, while several comparisons yielded small effects. The later European Jev supplement yielded five interactions below the adopted absolute reference of 0.05; its distinct rubric and interrupted collection limit comparison with the chat systems. Some written evaluations explicitly invoked a mismatch between a source and its attributed position. Taken together, the numerical and verbal evidence supports source-position coherence as a plausible explanation, alongside competing accounts involving credibility, authenticity and interpretation of the task. The paper develops this inference through controlled comparisons, reports conditional post hoc p-values in an appendix, and documents the human decisions and delegated checks behind an AI-conducted study.

Comments28 pages, 6 figures, 2 tables. Preregistrations, materials and code available on GitHub. Preprint; not yet peer reviewed

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑