Talk, Evaluate, Diagnose: User-aware Agent Evaluation with Automated Error Analysis
谈话、评估、诊断:基于用户意识的代理评估与自动错误分析
机构 * SAP ; Stanford University(斯坦福大学)
AI总结 本文提出TED框架,通过用户角色模板、自动评分和错误分析,提升代理评估的全面性与效率,实验显示在模型和用户水平上取得8-10%的性能提升。
Comments Accepted as a conference paper at ICLR 2026. Code and dataset are available in the repository https://github.com/SAP-samples/agent-quality-inspect