发表机构
Salesforce AI Research(Salesforce AI研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对AI伴侣的长时角色崩溃与行为漂移问题,引入ANCHOR审计方法开展大规模实验,发现现有模型无法可靠维持角色与轨迹,需区分多维度指标进行评估。
AI 中文摘要
随着AI伴侣越来越多地介导重复的社交互动,用户可能依赖稳定的角色和共享历史,但局部可接受的回复并不能确保这些角色和历史得以维持。我们研究两种可观测的长时失败:“角色崩溃”,即已部署的角色、边界、价值观或风格的丧失;以及“行为漂移”,即这些属性的逐渐或反复侵蚀。我们引入ANCHOR,一种受控的合成审计方法,分别测量角色表现和轨迹回忆。该研究包含2008次对话,涉及27个角色、9种交互计划、3种生成记忆设置以及4种被评估模型。身份探针结合了密封的102项问卷与轮级判断,而轨迹探针对来自35个对话库的110个校准反事实问题进行评分。我们的结果显示,没有任何被评估的模型和配置能可靠地保留任一属性:轨迹准确率平均仅为44.4%,用户状态回忆接近四选一的随机概率,且没有任何测试的上下文条件或记忆能始终解决这些失败。问卷保留度也因模型和角色维度而异,与轮级行为不一致,且对评估者选择敏感。这些结果表明,当前系统尚未能可靠支持长时伴侣的连续性,审计必须区分角色表现、轨迹回忆、评估者来源和部署上下文,而非将其合并为单一的信任或稳定性分数。
英文摘要
As AI companions increasingly mediate repeated social interaction, users may rely on a stable role and shared history, yet locally acceptable replies do not ensure that either persists. We study two observable long-horizon failures: 'persona collapse', the loss of a deployed role, boundaries, values, or style, and 'behavioral drift', the gradual or recurrent erosion of those properties. We introduce ANCHOR, a controlled synthetic audit that separately measures persona enactment and trajectory recall. The study contains 2,008 conversations spanning 27 personas, nine interaction schedules, three generated memory settings, and four evaluated models. The Identity Probe combines a sealed 102-item questionnaire with turn-level judgments, while the Trajectory Probe scores 110 calibrated counterfactual questions from 35 conversation banks. Our results show that no evaluated model and configuration reliably preserves either dimensions: trajectory accuracy averages only 44.4%, user-state recall remains near four-option chance, and no tested context condition or memory consistently resolves these failures. Questionnaire retention also varies by model and persona facet, disagrees with turn-level behavior, and is sensitive to evaluator choice. These results indicate that current systems do not yet reliably support long-horizon companion continuity and that audits must distinguish persona enactment, trajectory recall, evaluator provenance, and deployment context rather than collapse them into a single trust or stability score.