InterSocialBench:用于伴侣机器人社交行为的人类与LLM偏好基准测试
InterSocialBench: Benchmarking Human and LLM Preferences for Companion-Robot Social Behavior
- Xdream Robotics
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
InterSocialBench提出一个包含210个家庭场景和18种行为的基准,结合人类与LLM偏好,通过结构化流程和评估指标揭示模型与人类行为分布差异,支持社交行为选择评估。
AI中文摘要:
伴侣机器人在日常生活中常面临多种可行行为均可能合适、但不同人偏好不同回应的情况。我们提出了InterSocialBench,一个包含210个家庭场景和18种高层次行为的基准测试,汇集了100名人类参与者的判断以及7个大语言模型在16种人格条件下的23,520条回应。每条人类标注都保留了一个首选行为,同时附有明确合适与不合适的候选行为。结构化的构建流程涵盖了行为备选项、竞争性情境线索以及相关历史和未来任务。评估区分了首选选择一致性与明确拒绝,并使用按场景分组的数据划分来训练可预测模型。简单的频率基线和人格投票基线说明了这些目标。在测试的提示条件下,模型与人类的行为分布存在差异,且在匹配响应数量后多样性差距依然存在:人类在每个场景中平均有4.68个不同选择,而模型为2.06至3.46个。人类在场景层面的多数一致性为51.5%,这描述了分歧而非普遍预测上限。InterSocialBench支持评估社交行为选择,而无需用单一共识标签取代个体判断。
英文摘要:
Companion robots face everyday situations in which several feasible behaviors may be appropriate, yet different people prefer different responses. We introduce InterSocialBench, a benchmark of 210 domestic scenarios and 18 high-level behaviors, pairing judgments from 100 human participants with 23,520 responses from seven large language models under 16 personality conditions. Each human annotation preserves a preferred action alongside explicitly appropriate and inappropriate candidates. A structured construction pipeline covers behavioral alternatives, competing situational cues, and relevant history and future tasks. Evaluation distinguishes preferred-choice agreement from explicit rejection, using scenario-grouped splits for trainable predictors. Simple frequency and persona-voting baselines illustrate these objectives. Across the tested prompts, model and human behavior distributions differ, and the diversity gap remains after matching response counts: humans exhibit 4.68 distinct choices per scenario, compared with 2.06--3.46 for the models. Human scenario-level plurality agreement is 51.5%, describing disagreement rather than a universal prediction ceiling. InterSocialBench supports evaluating social behavior selection without replacing individual judgments with a single consensus label.