发表机构
Quis Lab(Quis 实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RealCompanion基准基于十段真实AI伴侣对话,揭示记忆需求稀少且难以检测,并比较了不同智能体系统重建人物画像的成本与性能。
AI 中文摘要
一个与人类交谈数月的伴侣应该逐渐理解对方。它应当记住对方说过的话,推断对方的身份,并知道过去的哪些信息与当前消息相关。测试这一点需要真实个人的记录,而此类记录是私密的,因此基准测试会生成人物和问题,并预先确定重要内容。我们发布了\ench,即与AI伴侣的十段真实关系:27,218条消息,时间跨度最长120天,以对话形式发布,并附带四个派生文件:个人资料、人物画像、聊天真实标签和问题集,每个文件都引用了其依据的消息。每个聊天标签都带有生成它的推理轨迹,并逐阶段对照对话进行核查。我们得出三项发现。第一,过去的信息很少被需要,且往往相隔很远。汇总指标会产生误导:一个近期窗口能为95.9%的探针找到所需消息,其中仅2.2%需要记忆;在自然频率下,提供记录证据所带来的收益中,96%来自不需要记忆的消息。第二,我们尝试的所有检测器都无法在真实消息上判断何时需要记忆,而基于相同历史生成的作者问题会泄露线索;将相同消息标记为记忆会使它们的使用率提高十到十四个百分点。第三,三个智能体系统以31倍的成本差异重建人物画像,但F1分数相同。
英文摘要
An AI companion that talks with someone for months should come to understand them. It should know who they are, remember what they said, and recognize when something from the past matters now. Testing this needs real conversations, but real conversations are private, so existing benchmarks use invented people and invented questions. We release RealCompanion, ten real relationships between people and an AI companion, with 27,218 messages over up to 120 days. For each person, we release the full conversation, a profile, a persona, chat test items, and question test items. Every label points to the messages that support it, and every chat label comes with the reasoning that produced it. The real data shows three things. First, people rarely refer back. Only 3.4% of their messages depend on something said earlier, and when one does, the earlier message is usually far away (a median of 2,157 messages back). Averages hide this. Looking at the most recent messages finds the needed one 95.9% of the time overall, but only 2.2% of the time when it is far back. Second, AI systems cannot tell when the past matters. The detectors we tested barely beat chance on real messages, and when the same earlier messages are labeled "memories" instead of "earlier messages", models bring up the past 10 to 14 percentage points more often, even when nothing from the past is needed. Third, AI systems read more into a person than the person revealed. Three agent systems rebuild each persona equally well (F1 0.71). They see the person, and then imagine more. Understanding a person depends on knowing when their past matters and where what they shared ends, and only real conversations can test it.