arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2604.25186cs.CVcs.CEcs.MM

FCMBench-Video:文档视频智能基准测试

FCMBench-Video: Benchmarking Document Video Intelligence

  • AI Lab, Qifu Technology(启赋科技AI实验室)
  • College of Future Information Technology, Fudan University(复旦大学未来信息学院)
  • School of Future Technology, South China University of Technology(华南理工大学未来技术学院)
  • Pazhou Lab(琶洲实验室)

机构由 AI 辅助整理,请以论文原文为准。

Runze Cui, Fangxin Shang, Yehui Yang, Qing Yang, Yanwu Xu, Tao Chen

更新

AI总结:

本文提出FCMBench-Video基准测试,用于评估文档视频理解中的文档感知、时间定位和证据推理能力,通过真实场景数据验证系统性能,揭示不同任务的敏感性和能力差异。

AI中文摘要:

文档理解是金融信贷审核、开卡和远程验证中的关键能力,其中决策准确性与证据可追溯性均至关重要。与静态文档图像相比,文档视频呈现时间冗余且顺序展开的证据流,需要跨帧证据整合,并保留与真实性敏感和反欺诈审核相关的采集过程提示。我们引入FCMBench-Video,一个用于文档视频智能的基准测试,评估文档感知、时间定位和证据推理能力。为了在隐私合规的前提下获得大规模真实数据,我们组织建设为原子级采集和组合工作流,记录可重用的单文档片段,应用受控降解,并组装长格式多文档视频。FCMBench-Video由495个原子视频组成1,200个长格式视频,配以11,322个专家标注的问题-答案实例,涵盖28种文档类型,时长在20秒至60秒之间,包含5,960个中文和5,362个英文实例。对九个最近的视频-大语言模型(Video-MLLMs)的评估显示,FCMBench-Video在系统和能力上提供了有意义的分离:计数是最敏感于时长的任务,跨文档验证和证据推理探测更高层次的证据整合,而视觉提示注入提供了互补的鲁棒性维度。总体得分分布广泛且近似钟形,表明该基准测试既不饱和也不被琐碎案例主导。这些结果使FCMBench-Video成为可重复使用的基准测试,用于跟踪视频-大语言模型在文档视频理解上的进展,并探测真实性敏感的信贷领域应用中的能力边界。

英文摘要:

Document understanding is a critical capability in financial credit review, onboarding, and remote verification, where both decision accuracy and evidence traceability matter. Compared with static document images, document videos present a temporally redundant and sequentially unfolding evidence stream, require evidence integration across frames, and preserve acquisition-process cues relevant to authenticity-sensitive and anti-fraud review. We introduce FCMBench-Video, a benchmark for document-video intelligence that evaluates document perception, temporal grounding, and evidence-grounded reasoning under realistic capture conditions. For privacy-compliant yet realistic data at scale, we organize construction as an atomic-acquisition and composition workflow that records reusable single-document clips, applies controlled degradations, and assembles long-form multi-document videos with prescribed temporal spans. FCMBench-Video is built from 495 atomic videos composed into 1,200 long-form videos paired with 11,322 expert-annotated question--answer instances, covering 28 document types over 20s--60s duration tiers and 5,960 Chinese / 5,362 English instances. Evaluations on nine recent Video-MLLMs show that FCMBench-Video provides meaningful separation across systems and capabilities: counting is the most duration-sensitive task, Cross-Document Validation and Evidence-Grounded Selection probe higher-level evidence integration, and Visual Prompt Injection provides a complementary robustness dimension. The overall score distribution is broad and approximately bell-shaped, indicating a benchmark that is neither saturated nor dominated by trivial cases. Together, these results position FCMBench-Video as a reproducible benchmark for tracking Video-MLLM progress on document-video understanding and probing capability boundaries in authenticity-sensitive credit-domain applications.

↑