arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

文本何时能替代视觉?图表推理中的结构性瓶颈

When Can Text Replace Vision? Structural Bottlenecks in Diagram Reasoning

Yunbei Zhang, Janet Wang, Jihun Hamm, Chandan K Reddy

arXiv 2609.39142首次发表:更新:

发表机构

Tulane University; Virginia Tech(杜兰大学; 弗吉尼亚理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出诊断协议,区分文本化图表推理错误源于信息缺失或利用不足,发现金标准结构显著优于视觉与学习文本,强调答案相关证据保留的重要性。

AI 中文摘要

结构化文本能否替代视觉进行图表推理?在文本化之后出现错误答案,可能是因为表示遗漏了问题所需的信息,也可能是因为求解器未能利用表示中已有的信息。我们引入了一种诊断协议来区分这些解释。使用相同的求解器模型和生成设置,我们比较三种输入条件:原始图像、由视觉语言模型提取的与问题无关的结构,或源自图表来源的金标准结构。有效性触发的恢复测试截断和模式失败,与问题相关的保真度衡量答案关键结构的保留程度,匹配边干预测试错误位置的影响。在保留的240个公共FlowGen图表测试集上,在冻结协议下评估,金标准结构达到87%的准确率,而直接视觉和学习文本均低于30%。汇总比较包括可能未在图像中打印的源派生关系标签,并使用不同的学习和金标准图编码,因此它不能单独隔离提取错误。仅重试无效提取使几乎所有公共表示模式有效,但准确率基本不变。公共学习文本相对于金标准的差距随着结构难度增加而增加一倍以上。与问题相关的拓扑比整个图拓扑更能预测正确性。在一项暴露干预研究中,单个答案相关边编辑将主要求解器的原始答案准确率降至接近零,而匹配的不相关编辑基本保持准确率。提供的结构比视觉需要更少的求解令牌,但学习获取在单次使用中消除了这一优势。这些比较激励我们根据保留的答案相关证据以及求解器使用该表示的能力来评估获取的文本。

英文摘要

Can structured text replace vision for diagram reasoning? A wrong answer after textualization can arise because the representation omits information the question needs, or because the solver fails to use information that is present. We introduce a diagnostic protocol to distinguish these explanations. Using the same solver model and generation settings, we compare three input conditions: the original image, question-blind structure extracted by a vision-language model, or gold structure derived from the diagram source. Validity-triggered recovery tests truncation and schema failure, question-relevant fidelity measures preservation of answer-critical structure, and matched edge interventions test the effect of error location. On a reserved holdout of 240 public FlowGen diagrams, evaluated under a frozen protocol, gold structure reaches 87% accuracy while direct vision and learned text both remain below 30%. The aggregate comparison includes source-derived relation labels that may not be printed in the image and uses different learned and gold graph encodings, so it does not isolate extraction error alone. Retrying only invalid extractions makes nearly every public representation schema-valid yet leaves accuracy essentially unchanged. The public learned-text deficit relative to gold more than doubles with structural difficulty. Question-relevant topology predicts correctness better than whole-graph topology. In an exposed intervention study, a single answer-relevant edge edit reduces the primary solver's original-answer accuracy to near zero, while matched irrelevant edits largely preserve it. Supplied structure requires fewer solving tokens than vision, but learned acquisition removes this advantage at single use. These comparisons motivate evaluating acquired text by the answer-relevant evidence it preserves and by the solver's ability to use that representation.

Comments33 pages, 16 figures. Code: https://github.com/yunbeizhang/text-for-vision

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑