arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多模态深度研究中忠实的图表生成:框架-证据协同自适应

Faithful Chart Generation for Multimodal Deep Research: Frame-Evidence Co-Adaptation

Yuxin Yue, Yingchen Zhang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke, Xueqi Cheng

arXiv 2610.00374首次发表:更新:

发表机构

Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences; University of Amsterdam(中国科学院计算技术研究所; 中国科学院大学; 阿姆斯特丹大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态深度研究中图表数值不忠实的问题,提出框架-证据协同自适应(FECA)方法,通过视觉框架与检索证据的迭代交互实现自适应规划,在100个真实主题上显著提升数值保真度并保持报告质量。

AI 中文摘要

多模态深度研究中的分析图表编码了定量声明,要求每个可视化的数值都忠实地基于支持性证据。与主要提供上下文信息的检索图像不同,图表要求数值保真度:可视化的数值不仅应在数量上与检索到的证据匹配,还应保留其原始含义和范围。然而,实现这种保真度仍然具有挑战性,因为当前系统通常在知道能从网络实际检索到哪些定量证据之前就构建可视化计划。因此,预定义的计划可能需要检索到的证据仅部分支持的实体、时间范围或比较维度。现有方法主要通过在图表计划固定后进行事后验证来解决此问题,这能够识别出不支持的数值,但使底层视觉框架保持不变。为了解决这一挑战,我们提出了框架-证据协同自适应(FECA),一种用于多模态深度研究的证据自适应可视化规划框架。受数据-框架理论中双向意义建构过程的启发,FECA将图表生成建模为视觉框架与检索证据之间的迭代交互。每个视觉框架都是自适应的:框架指导证据获取,而检索到的证据决定框架在渲染前是否应被接受、修订或丢弃。通过将可视化规划与证据可用性耦合,FECA将图表生成从固定计划验证转变为自适应的基于证据的视觉推理。在100个真实世界研究主题上的实验表明,FECA在保持报告质量和图表实用性的同时,显著提高了数值保真度。

英文摘要

Analytical charts in multimodal deep research encode quantitative claims, requiring every visualized value to be faithfully grounded in supporting evidence. Unlike retrieved images that mainly provide contextual information, charts require numerical fidelity: visualized values should not only match retrieved evidence quantitatively but also preserve its original meaning and scope. However, achieving such fidelity remains challenging because current systems usually construct visualization plans before knowing what quantitative evidence can actually be retrieved from the web. As a result, predefined plans may require entities, temporal ranges, or comparison dimensions that the retrieved evidence only partially supports. Existing approaches mainly address this issue through post-hoc verification after chart plans are fixed, enabling unsupported values to be identified but leaving the underlying visual frames unchanged. To address this challenge, we propose Frame-Evidence Co-Adaptation (FECA), an evidence-adaptive visual planning framework for multimodal deep research. Inspired by the bidirectional sensemaking process in Data-Frame Theory, FECA models chart generation as an iterative interaction between visual frames and retrieved evidence. Each visual frame is adaptive: the frame guides evidence acquisition, while retrieved evidence determines whether the frame should be accepted, revised, or dropped before rendering. By coupling visualization planning with evidence availability, FECA shifts chart generation from fixed-plan verification to adaptive evidence-grounded visual reasoning. Experiments on 100 real-world research topics show that FECA substantially improves numerical fidelity while preserving report quality and chart utility.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑