arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15455cs.RO

InterSocialBench:用于伴侣机器人社交行为的人类与LLM偏好基准测试

InterSocialBench: Benchmarking Human and LLM Preferences for Companion-Robot Social Behavior

  • Xdream Robotics

机构由 AI 辅助整理,请以论文原文为准。

Yaodan Xu, Boyang Guo, Yuqing Gu, Qingxin Zhang, Yiwen Deng, Meng Liu, Lintian Li

AI总结:

InterSocialBench提出一个包含210个家庭场景和18种行为的基准,结合人类与LLM偏好,通过结构化流程和评估指标揭示模型与人类行为分布差异,支持社交行为选择评估。

AI中文摘要:

伴侣机器人在日常生活中常面临多种可行行为均可能合适、但不同人偏好不同回应的情况。我们提出了InterSocialBench,一个包含210个家庭场景和18种高层次行为的基准测试,汇集了100名人类参与者的判断以及7个大语言模型在16种人格条件下的23,520条回应。每条人类标注都保留了一个首选行为,同时附有明确合适与不合适的候选行为。结构化的构建流程涵盖了行为备选项、竞争性情境线索以及相关历史和未来任务。评估区分了首选选择一致性与明确拒绝,并使用按场景分组的数据划分来训练可预测模型。简单的频率基线和人格投票基线说明了这些目标。在测试的提示条件下,模型与人类的行为分布存在差异,且在匹配响应数量后多样性差距依然存在:人类在每个场景中平均有4.68个不同选择,而模型为2.06至3.46个。人类在场景层面的多数一致性为51.5%,这描述了分歧而非普遍预测上限。InterSocialBench支持评估社交行为选择,而无需用单一共识标签取代个体判断。

英文摘要:

Companion robots face everyday situations in which several feasible behaviors may be appropriate, yet different people prefer different responses. We introduce InterSocialBench, a benchmark of 210 domestic scenarios and 18 high-level behaviors, pairing judgments from 100 human participants with 23,520 responses from seven large language models under 16 personality conditions. Each human annotation preserves a preferred action alongside explicitly appropriate and inappropriate candidates. A structured construction pipeline covers behavioral alternatives, competing situational cues, and relevant history and future tasks. Evaluation distinguishes preferred-choice agreement from explicit rejection, using scenario-grouped splits for trainable predictors. Simple frequency and persona-voting baselines illustrate these objectives. Across the tested prompts, model and human behavior distributions differ, and the diversity gap remains after matching response counts: humans exhibit 4.68 distinct choices per scenario, compared with 2.06--3.46 for the models. Human scenario-level plurality agreement is 51.5%, describing disagreement rather than a universal prediction ceiling. InterSocialBench supports evaluating social behavior selection without replacing individual judgments with a single consensus label.

补充信息

↑