AI 中文总结
本文提出Office Comprehension Bench,首个联合评估LLM对Word Excel和PowerPoint文件理解的基准,包含文件保真度问答和领域问答两个赛道,测试结构与视觉感知及多领域推理能力。
AI 中文摘要
我们介绍了Office Comprehension Bench (OCB),首个公开基准,用于联合评估LLM系统在Word Excel和PowerPoint文件及其变体上的理解能力。OCB包含两个赛道。文件保真度问答测试办公文档的结构和视觉感知,如表格、图表、嵌入图像、公式和特定应用元素。领域问答测试基于真实世界行业文档的专家级推理,涉及12个专业领域,需跨文档多步骤分析与综合。每个参考答案均分解为原子二元可评分声明,一组LLM裁判独立评分。即使最强前沿系统在默认推理模式下,领域问答仅达到约59.3%;在同一层级内增加推理深度对性能影响不大,而升级到更高产品层级则带来小幅提升。我们发布了数据集、评估工具、裁判提示和公开排行榜。
英文摘要
We introduce Office Comprehension Bench (OCB), the first public benchmark to jointly evaluate LLM systems on Word, Excel, and PowerPoint comprehension over native file formats (.docx, .xlsx, .pptx) and their variants. OCB consists of two tracks. File Fidelity Q&A tests structural and visual perception of office artifacts - tables, charts, embedded images, formulas, and app-specific elements such as headers, speaker notes, and named ranges. Domain Q&A tests expert-level reasoning grounded in real-world industry documents across 12 professional domains, with queries requiring multi-step analysis and synthesis across documents. Each reference answer is decomposed into atomic, binary-gradable claims, and an ensemble of LLM judges scores responses against each claim independently. Even the strongest frontier system in its default reasoning mode reaches only about 59.3% on Domain Q&A; increasing thinking depth within a tier does not move performance materially, while moving to a higher product tier yields modest gains. We release the dataset, evaluation tooling, judge prompt, and a public leaderboard.
CommentsAccepted at Findings of EMNLP 2026