证据所在之处:审计AI伴侣的自我描述
Where the Evidence Lives: Auditing AI Companions' Self-Descriptions
浏览论文内容
中文总结 AI 辅助
本文提出审计程序,将AI伴侣的自我描述与用户判断及实现记录对照,应用于Lita系统,发现记忆机制未运行,强调应报告无法证实的内容。
中文摘要 AI 辅助
伴侣智能体描述自己:它们会记忆,它们理解用户,关系已改变它们。我们认为,这类描述以及看似证实它们的体验评分,只有在证据属于用户的情况下才可被用户核查:即在智能体的行为中,或在用户自身中。当证据存在于机器内部时,流畅的自我描述和中等偏上的评分并不能证明其背后的机制确实运行了。我们展示了一种审计程序,将智能体的自我描述与其用户的判断及其实现记录进行对照,将每项声明报告为已支持、矛盾或未解决,并将其应用于Lita——一个我们构建并部署给九位同事使用一个月的主动型伴侣。参与者认可了风格性声明,对关系性声明未予认可,并对记忆评分在中等水平或以上,而三个记忆层中有两个从未执行其累积步骤。承载记忆的智能体应报告其自我描述无法确立的内容。
英文摘要
Companion agents describe themselves: they remember, they understand their users, the relationship has changed them. We argue that such accounts, and the experience ratings that seem to confirm them, are checkable by users only where the evidence is theirs: in the agent's behavior, or in themselves. Where the evidence lives in the machinery, fluent self-description and moderately positive ratings do not establish that the mechanisms behind them ran. We demonstrate an audit procedure that sets an agent's self-description against its users' judgements and its implementation records, reporting each claim as supported, contradicted, or unresolved, and apply it to Lita, a proactive companion we built and deployed for a month with nine colleagues. Participants endorsed stylistic claims, withheld endorsement from relational ones, and rated memory at or above midpoint, while two of three memory layers had never executed their accumulation step. Memory-bearing agents should report what their self-descriptions cannot establish.
发表机构
- Mamezo Co., Ltd.(豆藏株式会社)
- Rikkyo University(立教大学)
机构由 AI 辅助整理,请以论文原文为准。