arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

社会视听问答的推理:我们处于什么阶段?

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand?

Koen P. de Vries, Xavier Alameda-Pineda, Estefanía Talavera, Stéphane Lathuilière

arXiv 2608.13239首次发表:更新:

发表机构

Inria; Univ. Grenoble Alpes; CNRS; University of Twente(法国国家信息与自动化研究所; 格勒诺布尔大学; 法国国家科学研究中心; 特文特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对社会视听问答研究,发现IntentBench存在高噪声,Vanilla SFT基线性能优于现有推理方法,仅用文本模态即可实现与视频相当的性能,并发布了IntentBench-Prime等资源。

AI 中文摘要

针对视听社会理解训练多模态大语言模型(Multimodal Large Language Models, MLLMs)是实现具身社会智能的关键步骤。思维链(Chain-of-thought, CoT)推理已成为主流方法,HumanOmniV2及其IntentBench基准是重要参考。在此背景下,本文报告三项发现:第一,IntentBench存在高噪声:约7%的问题无效,约23%无需视频输入即可轻易回答,我们移除受影响问题并发布Intentbench-Prime;第二,当前推理方法成本高昂且效果不佳,简单的Vanilla SFT基线在三个基准上的表现与现有推理方法相当或更优,且成本仅为其一小部分,这使其成为评估新型微调技术的重要基线;第三,我们的分析显示,仅从文本模态即可学习到大量先验知识,使用文本描述替代视频时,性能与Vanilla SFT相当。这些意外发现揭示了当前多模态大语言模型在社会理解方面的局限性,IntentBench-Prime、Vanilla SFT模型及代码已公开。

英文摘要

Training Multimodal Large Language Models for audio-visual social understanding is a crucial step toward embodied social intelligence. Chain-of-thought (CoT) reasoning has become the dominant approach, with HumanOmniV2 and its IntentBench benchmark as a prominent reference point. In this context, we report three findings. First, IntentBench is highly noisy: $\sim$7% of questions are broken and $\sim$23% are trivially answerable without the video input. We remove the affected questions and release Intentbench-Prime. Second, current reasoning approaches are expensive and surprisingly ineffective. A simple Vanilla SFT baseline matches or outperforms existing reasoning methods across three benchmarks at a fraction of the cost, establishing it as an essential baseline for evaluating novel fine-tuning techniques. Third, our analysis reveals that substantial priors can be learned solely from the text modality and that using a textual caption instead of the video yields performance on par with Vanilla SFT. These surprising findings reveal the limitations of current MLLMs when it comes to social understanding. IntentBench-Prime, Vanilla SFT model, and code are publicly available.

CommentsAccepted at HCMIW ECCV workshop. Code available here: https://github.com/koenv759/VanillaSFT

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑