发表机构
Kuaishou GameMind Lab(快手GameMind实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对角色扮演评估的黑箱问题,提出TRACE Bench框架,通过清单分解与对话跟踪实现更透明的评估,在覆盖度、鲁棒性等方面优于现有基准,支持闭环演化。
AI 中文摘要
角色扮演评估不应仅给出单一分数,而应揭示哪些角色要求被测试、哪些未通过,以及哪些对话证据支持判断。我们提出TRACE Bench,一个任务驱动的智能体清单评估框架。它离线将每个角色档案分解为固定清单,再使用用户智能体与目标角色扮演模型自然对话,同时根据模型响应私下更新清单状态。因此分数可追溯至清单项和支持的对话轮次,而非黑箱整体印象。为进行覆盖交叉验证,我们对照同一角色衍生清单审核MiniMax角色扮演基准中发布的M2自由对话对话记录。发布的自由对话记录仅覆盖73.74%的关键角色档案要点,而TRACE Bench在更少轮次中达到99.91%的覆盖率。鲁棒性实验显示,在重复运行和用户智能体替换下排名稳定。在26个模型上,TRACE Bench报告了整体排名及能力细分和清单轨迹。它还支持闭环基准演化,从失败轨迹中提炼经证实有效的验证方法,使后续评估能更可靠地引出和检查观察到的失败模式。
英文摘要
Roleplay evaluation should do more than assign a single score: it should reveal which role requirements were tested, which failed, and which dialogue evidence supports the judgment. We propose TRACE Bench, a task-driven agentic checklist evaluation framework. It decomposes each role profile offline into a fixed checklist, then uses a User Agent to converse naturally with the target roleplay model while privately updating checklist states from model responses. Scores therefore trace back to checklist items and supporting dialogue turns rather than a black-box holistic impression. For coverage cross-validation, we audit released M2 free-dialogue transcripts from the MiniMax Role-play Benchmark against the same role-derived checklist. The released free-chat transcripts cover only 73.74% of key role-profile points, whereas TRACE Bench reaches 99.91% coverage in fewer turns. Robustness experiments show stable rankings under repeated runs and User Agent replacement. Across 26 models, TRACE Bench reports overall rankings together with capability breakdowns and checklist traces. It also supports Closed-Loop Benchmark Evolution, distilling verification methods proven effective in failed traces so later evaluations can more reliably elicit and examine observed failure modes.
CommentsProject page: https://kuaishou-gamemind.github.io/projects/trace_bench/. Code: https://github.com/KuaishouGameMind/TRACE-Bench