FIRM-Video:可靠文本到视频奖励建模的评分前检查机制
FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling
- Shanghai Jiao Tong University(上海交通大学)
- Tencent Youtu Lab(腾讯优图实验室)
- Tongji University(同济大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出基于“评分前检查”的FIRM-Video框架,构建相关数据集与基准,其衍生模型在文本到视频奖励建模相关任务上取得最优表现。
AI中文摘要:
可靠的奖励模型对于文本到视频的评估与对齐至关重要,但评估精度与推理效率之间的权衡对训练监督质量提出了高要求。现有方法常依赖带有固定规则的整体评判或开放式推理,导致检查不完整、解释不忠实及归因混乱。本文提出FIRM-Video,一种基于“评分前检查”原则的统一清单驱动数据构建框架:构建维度特定的清单,针对时间视觉证据验证各准则,仅聚合经验证的决策。对于指令遵循任务,FIRM-Video将提示分解为带权重的原子需求;对于世界一致性任务,构建经提示校准、针对目标的检查,基于可见实体与动作;对于感知质量任务,采用通用的视觉缺陷分类法。经验证的准则与分数进一步转换为自然语言分析,用于端到端奖励建模。随后,本文构建包含29348个视频的88044个维度特定实例的FIRM-Video-90K,并引入包含250个视频的750个逐点人工标注的FIRM-Video-Bench。基于Qwen3-VL的FIRM-Video-8B在FIRM-Video-Bench上取得最佳整体MAE,且在3种视频生成器的Best-of-8采样中,持续获得最高的VBench总分、质量分与语义分。
英文摘要:
Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evaluation accuracy and inference efficiency places high demands on the quality of training supervision. Existing approaches often rely on holistic judges with fixed rubrics or open-ended reasoning, leading to incomplete inspection, unfaithful justification, and entangled attribution. We introduce FIRM-Video, a unified checklist-driven data construction framework based on a check-before-score principle: construct dimension-specific checklists, verify each criterion against temporal visual evidence, and aggregate only verified decisions. For Instruction Following, FIRM-Video decomposes prompts into weighted atomic requirements; for World Coherence, it constructs prompt-calibrated, target-specific checks grounded in visible entities and actions; and for Perceptual Quality, it applies a generic taxonomy of visual defects. The verified criteria and scores are further transformed into natural-language analyses for end-to-end reward modeling. Subsequently, we construct FIRM-Video-90K with 88,044 dimension-specific instances from 29,348 videos, and introduce FIRM-Video-Bench with 750 point-wise human annotations across 250 videos. The Qwen3-VL-based FIRM-Video-8B achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.