arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MMHBench:用于长视频心理健康理解的多视角基准测试集

MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos

Jinpeng Hu, Erqiang Wang, Shan Wang, Zhuo Li, Peipei Song, Xun Yang, Meng Wang

arXiv 2607.27895首次发表:更新:

AI 中文总结

本文提出MMHBench多视角心理健康理解基准,含268个长视频与2184条问题,采用MAQG框架生成问题,评估22个多模态大语言模型后发现该任务仍具挑战性。

AI 中文摘要

长视频中的心理健康理解需要对可观察行为、人际语境和潜在心理状态进行细致推理。现有基准测试集大多将该任务简化为粗粒度分类,难以判断模型是否真正理解心理现象,还是依赖表面关联。为解决这一局限,本文提出MMHBench,一个用于多视角心理健康理解的综合多模态基准,包含268个长视频和2184条精心策划的问题。MMHBench将评估分为两个互补设置:(1)第三人称评估,含605条问题,聚焦可观察行为和多模态证据的解读;(2)第一人称视角采择,含1579条问题,需基于视角条件推理,识别多模态证据支持的心理状态解读。本文提出多智能体问题生成(MAQG)框架,模拟不同社会角色生成多视角问题,生成的问题经多角色反馈与迭代优化,再经专家指导验证以确保质量与有效性。对22个代表性多模态大语言模型(MLLMs,涵盖开源及领先闭源模型)的广泛评估表明,长视频心理健康理解仍极具挑战性。

英文摘要

Mental health understanding in long-form videos requires nuanced reasoning over observable behavior, interpersonal context, and latent psychological states. Existing benchmarks largely reduce this task to coarse-grained classification, providing limited insight into whether models truly understand psychological phenomena or rely on superficial correlations. To address this limitation, we introduce MMHBench, a comprehensive multimodal benchmark for multi-perspective mental health understanding, comprising 268 long-form videos and 2,184 carefully curated questions. MMHBench organizes the evaluation into two complementary settings: (1) third-person assessment, consisting of 605 questions that focus on the interpretation of observable behaviors and multimodal evidence, and (2) first-person perspective-taking, comprising 1,579 questions that require perspective-conditioned reasoning to identify the interpretation of the mental state supported by the available multimodal evidence. We propose a Multi-Agent Question Generation (MAQG) framework that simulates diverse social roles to synthesize questions from multiple perspectives. The generated questions are refined through multi-role feedback and iterative optimization, followed by expert-guided verification to ensure quality and validity. Extensive evaluation of 22 representative multimodal large language models (MLLMs), spanning both open-source and leading closed-source models, demonstrates that long-form video mental health understanding remains highly challenging.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑