发表机构
South China University of Technology(华南理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对真实场景退化影响多模态大语言模型文档问答的问题,提出无需训练的DocIntent框架,通过评估可回答性、选择性调用恢复工具及比较式回滚机制,提升了各类MLLM在WildDoc基准上的表现。
AI 中文摘要
模糊、阴影、畸变、莫尔纹等真实场景退化严重影响多模态大语言模型(MLLM)的文档问答能力。在视觉问答(VQA)前应用恢复工具是直观解决方案,但现有恢复方法仍存在局限:手动设计并执行恢复策略既费力又需领域专业知识。智能体式恢复为自动化提供了新可能,不过现有框架主要针对自然图像,且追求感知质量,忽略了恢复应服务于下游任务而非优化通用图像质量指标。为此,我们探索智能体式恢复在真实场景退化文档VQA中的价值,提出DocIntent,这是一个无需训练的可回答性引导智能体式恢复框架。DocIntent首先评估问题的可回答性,再识别与任务相关的退化并选择性调用恢复工具;其采用基于比较的回滚机制验证每一步恢复操作,当与问题相关的证据可辨识度降低时则回滚该操作。整个过程无需额外的预训练退化分类器或图像质量评估模型。在WildDoc基准上的大量实验表明,DocIntent能持续提升各类开源及闭源MLLM的平均得分与一致性,代码及实验数据将公开提供。
英文摘要
Real-world degradations such as blur, shadow, distortion, and moire patterns severely impair the document question-answering capabilities of Multimodal Large Language Models (MLLMs). Applying restoration tools before Visual Question Answering (VQA) is an intuitive solution. However, existing restoration approaches remain limited, as manually designing and executing restoration strategies is labor-intensive and requires domain expertise. Agentic restoration offers new possibilities for automation, yet existing frameworks primarily target natural images and pursue perceptual quality, overlooking that restoration should serve downstream tasks rather than optimize generic image quality metrics. To this end, we explore the value of agentic restoration for real-world degraded document VQA and propose DocIntent, a training-free Answerability-Guided Agentic Restoration framework. DocIntent first assesses question answerability, then identifies task-relevant degradations and selectively invokes restoration tools. A Comparison-Based Rollback mechanism validates each restoration step and reverts it when question-relevant evidence becomes less decipherable. The entire process requires no additional pretrained degradation classifier or image quality assessment model. Extensive experiments on the WildDoc benchmark show that DocIntent consistently improves the average score and consistency of different open- and closed-source MLLMs. The code and experimental data will be publicly available.
Comments25 pages, 14 figures