FinixDoc:重新思考超出饱和基准的金融文档解析
FinixDoc: Rethinking Financial Document Parsing Beyond Saturated Benchmarks
- Ant Group(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对金融文档解析基准与实际需求的差距,提出端到端智能体解析系统FinixDoc,核心为4B规模视觉语言模型FinixDoc-VL,在自建评估套件FinixDocBench上实现最优性能。
AI中文摘要:
金融文档解析需要当前基准通常无法反映的准确性、结构一致性和可验证性。我们提出FinixDoc,这是一个针对真实世界金融文档的端到端智能体解析系统,其核心解析器为基于Qwen3-VL-4B构建的40亿参数规模视觉语言模型FinixDoc-VL。为刻画基准性能与部署性能之间的差距,我们引入了沿两个实用轴(视觉质量和文档规模)组织的文档解析能力矩阵。在该矩阵的指导下,FinixDoc-VL采用结合了同形字感知对比学习和带复合领域特定奖励的多阶段强化学习的领域适配方案进行训练。为更好地利用我们在低质量金融文档数据上积累的优势并支持大规模高质量数据生产,我们进一步构建了带置信度感知专家审核的人在环Data Factory流水线。在评估环节,我们构建了覆盖数字原生、相机拍摄、超大页面及内部工作流场景的金融领域评估套件FinixDocBench,同时发布了经合规审核的子集。在其主要子集上,FinixDoc-VL在评估基准中取得最高综合得分(81.43),领先次优开源模型5.13分,在内部金融工作流(FinixInner)上的提升最大(84.08对78.73)。
英文摘要:
Financial document parsing requires accuracy, structural consistency, and verifiability that current benchmarks often fail to reflect. We present FinixDoc, an end-to-end agentic parsing system for real-world financial documents, with FinixDoc-VL, a 4B-scale vision-language model built on Qwen3-VL-4B, as its core parser. To characterize the gap between benchmark and deployment performance, we introduce a Document Parsing Capability Matrix organized along two practical axes: visual quality and document scale. Guided by this matrix, FinixDoc-VL is trained with a domain-adapted recipe combining homoglyph-aware contrastive learning and multi-stage reinforcement learning with composite domain-specific rewards. To better leverage our accumulated advantage in low-quality financial-document data and support large-scale, high-quality data production, we further build a human-in-the-loop Data Factory pipeline with confidence-aware expert review. For evaluation, we construct FinixDocBench, a financial-domain evaluation suite covering digital-native, camera-captured, ultra-large-page, and internal-workflow scenarios, with a compliance-reviewed subset released alongside this technical report. On its main subsets, FinixDoc-VL achieves the highest overall score (81.43) among evaluated baselines, outperforming the next-best open-source model by 5.13 points, with the largest gains on internal financial workflows (FinixInner: 84.08 vs. 78.73).