发表机构
University of Nebraska at Omaha; San Jose State University(内布拉斯加大学奥马哈分校; 圣何塞州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究以《反人类牌》风格任务开展实验,发现大语言模型可通过裁判的先前选择及理由,提升近似另一模型幽默偏好的准确率,为跨模型偏好建模提供操作层面的证据。
AI 中文摘要
本文在受控的《反人类牌》(Cards Against Humanity)风格任务中,研究一个大语言模型能否近似另一个大语言模型的幽默偏好。实验设置二元幽默选择任务,确保成功无法源于自我偏好,评估两个模型:作为“裁判(Czar)”的GPT-4o和作为“玩家(Player)”的Claude Opus-4.5。通过反射单元稳定性程序分离出244组手牌,两个模型对这些手牌持有确定但相反的偏好,分为97组的上下文池和147组的保留测试池。随后在五个梯度条件下评估玩家的表现:默认自我偏好、通用裁判建模指令、模型识别的裁判、裁判的先前选择、带理由的裁判先前选择。该梯度设计用于分离两类改进来源:框架效应(仅告知玩家关注裁判,不提供裁判的任何行为)和直接行为证据(向玩家展示裁判的先前选择)。玩家准确率在条件1中为0.7%,仅框架条件下升至19.0%和25.9%,提供行为证据和理由后进一步升至72.8%和82.3%。综合Cochran Q检验和成对McNemar检验证实,梯度中的每一步都产生显著改进。结果表明,角色指令和模型身份仅产生微小增益,而行为证据(尤其伴随理由时)支持显著的跨模型偏好建模。研究结果被解释为类似心理理论的行为,属于操作层面而非表征层面:玩家从自我偏好转向另一智能体的已证明偏好,不涉及对心理状态底层表征的任何主张。
英文摘要
This paper investigates whether one large language model can approximate the humor preferences of another in a controlled Cards Against Humanity-style task. Two models - GPT-4o as Czar and Claude Opus-4.5 as Player - are evaluated on a binary humor-selection task constructed so that success cannot follow from self-preference. A reflected-cell stability procedure isolates 244 hands on which the two models hold deterministic but opposite preferences, partitioned into a 97-hand context pool and a 147-hand held-out test pool. The Player is then evaluated across five graded conditions: default self-preference, generic Czar-modeling instruction, model-identified Czar, prior Czar selections, and prior Czar selections with rationales. This gradient is designed to separate two sources of improvement: framing effects, in which the Player is told to attend to a Czar without seeing any of the Czar's behavior, and direct behavioral evidence, in which the Player is shown the Czar's prior choices. Player accuracy increased from 0.7% in Condition 1 to 19.0% and 25.9% in the framing-only conditions, and then rose to 72.8% and 82.3% once behavioral evidence and rationales were provided. An omnibus Cochran's Q test and pairwise McNemar tests confirmed that each step in the gradient produced a significant improvement. The results indicate that role instruction and model identity yield only modest gains, while behavioral evidence - especially when accompanied by rationales - supports substantial cross-model preference modeling. The findings are interpreted as theory-of-mind-like behavior in an operational rather than representational sense: the Player shifts away from self-preference toward another agent's demonstrated preferences, without any claim about an underlying representation of mental states.
Comments11 pages, 2 figures, 1 table