PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks
PPT-Eval: 面向PowerPoint任务的计算机使用智能体基准
机构 * Carnegie Mellon University(卡内基梅隆大学) ; Microsoft(微软) ; UMass Amherst(马萨诸塞大学阿默斯特分校) ; Google(谷歌) ; Snowflake
AI总结 提出PPT-Eval基准,包含120个PowerPoint任务,并设计基于评分标准的评估框架,通过部分评分和自然语言反馈衡量智能体表现,发现现有最强模型仅达45%成功率。
Comments Proceedings of the 43rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026