AI 中文总结
研究多模态文档问答中高效训练后优化问题,提出感知 - RFT 框架,用组相对策略优化绕过推理令牌直接对齐视觉与基础输出,通过构建变体评估推理必要性,发现启用推理模型有优势,还识别基础差异,表明早期转换可减少训练数据并保持精度。
AI 中文摘要
高效的多模态文档问答以及明确的视觉基础,即定位支持每个答案的精确文档区域,仍然是一个开放的挑战。当前方法分为监督微调(SFT)和以推理为中心的强化学习(RL)。SFT 需要大量标注数据集且达到优化平台期,RL 依赖冗长的中间轨迹且增加推理令牌成本却无明显益处。我们引入感知 - RFT,一种将组相对策略优化(GRPO)应用于多模态文档问答的训练框架,绕过中间推理令牌直接将视觉特征与结构化基础输出对齐。为严格评估推理的必要性,我们构建相同奖励设置下的推理变体。发现启用推理的模型在训练期间抑制推理轨迹,在 4B 参数规模下收敛到基于直接感知的策略,将每个查询的推理令牌长度减少 60%以上,而启用推理的 RL 表现不如仅感知训练。通过对 Qwen3 - VL - 4B 优化动态的细粒度分析,确认文本域训练后建立的 SFT 饱和和冷启动 RL 不稳定性扩展到多模态,并识别出先前未表征的基础差异:在联合 RL 优化下,在两个分布外(OOD)基准(4828 个样本)上语义鲁棒性和几何精度之间的选择性权衡。还表明早期 SFT→RL 转换以少 65%的训练数据实现了可比的精度。
英文摘要
Efficient multimodal document question answering with explicit visual grounding, locating the precise document region that supports each answer remains an open challenge. Current approaches bifurcate into Supervised Fine-Tuning (SFT), which requires large annotated datasets and reaches optimization plateaus, and reasoning-centric Reinforcement Learning (RL), which depends on verbose intermediate traces that inflate inference token cost without clear benefit. We introduce Perception-RFT, a training framework that applies Group Relative Policy Optimization (GRPO) to multimodal document QA, bypassing intermediate reasoning tokens to directly align visual features with structured grounding outputs. To rigorously evaluate the necessity of reasoning, we construct a reasoning variant under identical reward settings. We find that reasoning-enabled models suppress their reasoning traces during training, converging to direct perception-based policies at the 4B parameter scale, reducing per-query inference token length by more than 60%, while reasoning-enabled RL underperforms perception-only training. Through a fine-grained analysis of Qwen3-VL-4B optimization dynamics, we confirm that SFT saturation and cold-start RL instability established in text-domain post-training extend to multimodal, and identify a previously uncharacterized Grounding Divergence: a selective trade-off between semantic robustness and geometric precision on two out of distribution (OOD) benchmarks (4,828 samples) under joint RL optimization. We further show that an early SFT$\rightarrow$RL transition achieves comparable precision with 65% less training data.
CommentsAccepted at ICML 2026, Workshop on Efficient Multimodal Question Answering (EMM-QA)