CompanionBench:一个基于理论、扎根于真实世界的AI情感陪伴基准
CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship
浏览论文内容
中文总结 AI 辅助
该研究推出基于理论与真实数据的交互式双语AI情感陪伴基准CompanionBench,评估28个智能体发现其能力差异,指出角色扮演智能体表现差、核心能力存弱点,将发布500对中英平行数据及评估代码。
中文摘要 AI 辅助
大型语言模型(LLM)陪伴系统正大规模部署于对个人至关重要的场景中,却缺乏完善的评估机制。现有基准使用人工编写的场景和提示式模拟器,将同理心汇总为单一分数,且忽视了评判者偏差,如同家族偏好和量表漂移。我们推出CompanionBench,一个交互式双语基准。据我们所知,它是首个将场景和训练后的用户模拟器都基于去标识化真实世界数据的陪伴基准。一个隐藏的披露门控根据智能体自身行为分支每个角色的轨迹,在不编写对话脚本的情况下控制交互状态空间。我们将心理学和咨询领域25种理论衍生的10种能力进行操作化,其中4种是先前工作未明确评分的:容纳模糊性、自体客体响应性、积极共振和校准挑战。智能体在两个互补维度上接受评估:主观的10项能力 rubric(评分标准)和衡量是否获得更深层次披露的确定性指标。跨家族评审小组削弱了同家族偏好;项目反应理论模型将智能体质量与评判者严格程度分离。理论确定了测量内容和角色结构;真实数据提供事件、历史和概况——理论提供覆盖范围,数据提供真实性。排名在两种语言中均可复现(ZH语言的rho=0.996,EN语言的rho=0.953)。对28个智能体的评估揭示了被汇总分数掩盖的能力层面差异:情绪调节和校准挑战仍是常见弱点;容纳模糊性的区分度最高。角色扮演智能体排名接近末尾:沉浸感并不意味着关系能力。在所有智能体中,主要失败模式是用表面温暖替代实质性关系支持。我们将发布500对中英平行数据及评估代码。
英文摘要
LLM companions are deployed at scale in personally consequential settings, yet poorly evaluated. Existing benchmarks use hand-authored scenarios and prompted simulators, aggregate empathy into one score, and overlook judge biases such as same-family favoritism and scale drift. We introduce CompanionBench, an interactive bilingual benchmark. To our knowledge, it is the first companion benchmark to ground both its scenarios and a trained user simulator in de-identified real-world data. A hidden disclosure gate branches each persona's trajectory on the agent's own behavior, controlling the interaction state space without scripting dialogue. We operationalize ten capabilities derived from 25 theories across psychology and counseling, four of them not graded explicitly by prior work: holding ambiguity, selfobject responsiveness, positive resonance and calibrated challenge. Agents are assessed on two complementary axes: a subjective ten-capability rubric and a deterministic measure of whether deeper disclosure was earned. A cross-family panel dilutes same-family favoritism; an Item Response Theory model separates agent quality from judge severity. Theory fixes what to measure and how personas are structured; real data supply events, history, and profiles -- coverage from theory, authenticity from data. Rankings are reproducible in both languages (rho = 0.996 ZH / 0.953 EN). Evaluating 28 agents reveals capability-level differences obscured by aggregate scores. Emotion regulation and calibrated challenge remain common weaknesses; holding ambiguity discriminates most. Role-play agents rank near the bottom: immersion does not imply relational competence. Across agents, the dominant failure mode is substituting surface warmth for substantive relational support. We will release 500 Chinese-English parallel pairs and the evaluation code.