arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Video-FLAIR:不是是否推理,而是如何推理

Video-FLAIR: Not Whether to Reason, But How

Yogesh Kulkarni, Pooyan Fazli

arXiv 2608.26495首次发表:更新:

发表机构

Arizona State University(亚利桑那州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出 Video-FLAIR 训练框架,通过强化学习让模型为多模态查询选择适配的推理模式,在多个视频多模态基准上提升准确率并降低 token 用量。

AI 中文摘要

多模态查询可能需要不同类型的推理:部分查询可通过感知推理直接从视觉信号提取信息,部分则需要结合观察结果的组合推理,或评估竞争假设的 deliberative 推理(审议式推理)。然而,现有许多方法对所有查询采用统一推理策略,导致简单任务出现不必要的计算,复杂任务推理不足。我们提出 Video-FLAIR,这是一个训练框架,通过强化学习为每个查询学习选择合适的推理模式。训练时,模型针对同一提示生成三种模式下的响应,便于直接比较;复合奖励会对比这些响应,依据正确性、 grounding( grounding 指与输入的匹配度)和成本选出最有效的响应,同时抑制无依据或错位的审议,从而在无需每个查询标注的情况下,为自适应推理提供监督信号。Video-FLAIR 在 MathVista 上较 Qwen2.5-VL 基础模型准确率提升 5.4,在 Video-Holmes 和 Video-MMMU 上各提升 4.8,同时将平均 token 用量从始终推理基线的 417 降至 95。

英文摘要

Multimodal queries can require different types of reasoning. Some can be answered via perceptual reasoning, extracting information directly from the visual signal, while others require compositional reasoning that combines observations or deliberative reasoning that evaluates competing hypotheses. However, many existing methods apply a uniform reasoning strategy across queries, leading to unnecessary computation on simple tasks and insufficient reasoning on complex ones. We introduce Video-FLAIR, a training framework that learns to select the appropriate reasoning mode for each query using reinforcement learning. During training, the model generates responses under all three modes for the same prompt, enabling direct comparison. A composite reward compares these responses to favor the most effective one based on correctness, grounding, and cost, while discouraging unsupported or misaligned deliberation. This yields a supervision signal for learning adaptive reasoning without per-query annotations. Video-FLAIR improves accuracy over the Qwen2.5-VL base model by +5.4 on MathVista, +4.8 on Video-Holmes, and +4.8 on Video-MMMU, while reducing average token usage to 95 compared to 417 for always-thinking baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑