arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越编辑画布:OOXML导入大语言模型过程中的证据分歧

Beyond the Editing Canvas: Evidence Divergence in OOXML-to-LLM Ingestion

Side Liu, Jiangpeng Liu, Jinwen Xin, Guojun Peng, Jiang Ming

arXiv 2608.25880首次发表:更新:

AI 中文总结

本文发现OOXML导入LLM流程存在证据分叉,同一份OOXML文件在Office与LLM中呈现不同视图,21种分叉经测试,多数接口会返回Office未显示的陷阱,开源项目默认导入路径集中于受影响提取器。

AI 中文摘要

大语言模型(LLM)流程越来越多地将Office Open XML(OOXML,即Word、Excel和PowerPoint文件)作为金融、合规及检索增强工作流中的一级证据,默认假设其语义完整性——即模型所使用的证据与Microsoft Office套件编辑画布中显示的内容一致。本文表明,这一假设在OOXML导入LLM的流程中可能不成立:同一份符合规范的OOXML文件在Microsoft Office中会呈现一种证据视图,而在为LLM提取时会呈现另一种,每种视图都被其使用者视为权威,我们将这种情况称为“多元真实值”。证据导入契约很少明确说明哪种视图及语义角色会成为模型证据,或保留该证据的推导方式,我们将引发此类分歧的、基于规范的OOXML结构称为“证据分叉”。我们系统遍历并挖掘OOXML规范,在Excel、Word和PowerPoint中确认了21种证据分叉,涵盖视图构建的六个维度;我们的提取面板中的13种工具均至少在一种分叉中产生证据。我们测试了四个原生导入LLM API和七个网络聊天机器人,每份测试文档都包含一个陷阱:一个被提取暴露但未在Office中显示的、与任务相关的事实。在针对21种机制的评估中,四个API在48%-76%的试验中返回该陷阱;在21种机制中的20种里,十一个接口中至少有一个返回了该陷阱。我们的测量进一步表明,暴露情况由上游的导入路径和提取器配置决定;对十六个流行开源LLM项目的源级调查还显示,默认OOXML导入路径集中在受影响的提取器家族中。

英文摘要

LLM pipelines increasingly ingest Office Open XML (OOXML) documents (Word, Excel, and PowerPoint files) as first-class evidence in financial, compliance, and retrieval-augmented workflows, implicitly assuming semantic integrity: that the evidence consumed by the model matches the content shown in the Microsoft Office suite editing canvas. We show that this assumption can fail in OOXML-to-LLM pipelines. The same specification-valid OOXML file can yield one evidentiary view in Microsoft Office and another when extracted for an LLM. Each view is treated as authoritative by its consumer, a condition we call plural ground truth. The ingestion contract rarely states which view and semantic roles become model evidence or preserves how that evidence was derived. We call the specification-grounded OOXML constructions that induce such divergence evidence forks. We systematically traverse and mine the OOXML specification and confirm 21 evidence forks across Excel, Word, and PowerPoint, spanning six dimensions of view construction. All 13 tools in our extraction panel emit evidence from at least one fork. We test four native-ingestion LLM APIs and seven web chatbots. Each test document carries a trap: a task-relevant fact exposed by extraction but not shown in Office. Across this 21-mechanism evaluation, the four APIs return the trap in 48--76% of trials. For 20 of 21 mechanisms, at least one of the eleven interfaces returns the trap. Our measurements further show that exposure is shaped upstream of the model by the ingestion path and extractor configuration. A source-level survey of sixteen popular open-source LLM projects further shows that default OOXML ingestion paths concentrate on affected extractor families.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑