arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FriendBench:人类与多模态大语言模型的二元熟悉度推理基准

FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models

Jeffrey M. Girard, Jason Z. Zheng, Jacqueline R. Vertino, Antony D'Avirro, Benjamin Peloquin

arXiv 2607.29602首次发表:更新:

发表机构

Fluid Concepts Research(流体概念研究机构)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出FriendBench基准,通过20秒二元破冰对话片段推断两人熟悉度,对比7家公司26个模型与人类的表现,发现模型与人类准确率无统计差异但先验倾向不同,且仅人类能从可见行为中获益,同时发布了相关资源。

AI 中文摘要

解读社交情境往往依赖行为而非仅言语。我们推出FriendBench,这是一个用于从20秒二元破冰对话片段中推断两人是已熟悉还是初次见面的基准。每对都回答相同类型的提示,因此仅互动方式能揭示答案。在文本、音频和视频模态下,我们将来自7家公司的26个模型与96组平衡二元对的匹配人类面板进行比较。最佳模型与人类群体在各模态的准确率上无统计差异,但达成路径不同:人类在两种答案间保持平衡,而最强模型倾向于“陌生人”——这是有效先验的差异,而非判别能力的差异。更丰富的模态对两者的帮助不均等,且仅人类能从语音之外的可见行为中获益。我们发布了刺激材料、人类评分和模型预测结果。

英文摘要

Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic ice-breaker conversation. Every pair answers the same type of prompt, so only the manner of interaction can reveal the answer. Across text, audio, and video, we compare 26 models from seven companies against matched human panels over 96 balanced dyads. The best model and the human crowd are statistically indistinguishable on accuracy in every modality, but reach it differently: humans stay balanced across the two answers, while the strongest models favor ``stranger.'' This is a difference in effective prior, not in discrimination. Richer channels help both unequally, and only humans gain from visible behavior on top of speech. We release the stimuli, human ratings, and model predictions.

Comments17 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑