arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MIS-Bench:多模态大语言模型在心理治疗人际技能评估中的基准测试

MIS-Bench: Benchmarking Multimodal LLMs for Psychotherapeutic Interpersonal Skills Assessment

Yuhan Lu, Yi Yao, Hua Shen, Katie Aafjes-van Doorn, Zhaonan Wang

arXiv 2609.22778首次发表:更新:

发表机构

NYU Shanghai, New York University(上海纽约大学,纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出MIS-Bench基准,评估多模态大语言模型在心理治疗人际技能评估中的表现,发现其与专家一致性有限,并引入MIS-RAFT微调方法以提升评分准确性。

AI 中文摘要

多模态大语言模型(MLLMs)越来越多地被用作评估者,但它们在需要专家判断的专业评估任务中的可靠性仍不明确。我们在此背景下研究了心理治疗人际技能的评估问题,并引入了MIS-Bench,一个多模态人际技能(MIS)基准,包含996个心理治疗反应视频,这些视频在促进性人际技能的8个维度上进行了标注。在9个MLLM上,通过多种模态和提示设置,我们发现当前模型与人类专家的一致性仅为中等水平,多模态输入的收益不一致,且基于推理的提示带来的益处有限。为缩小这一差距,我们提出了MIS-RAFT,一种受RAFT启发、针对一位小数精度细粒度人际技能评分定制的回归感知微调方法。MIS-RAFT解决了自回归词元预测与标量值专家评估之间的不匹配问题,显著提高了与人类评分的一致性。总体而言,MIS-Bench揭示了通用多模态能力与专家级人际判断之间的明显差距,而MIS-RAFT为更可靠的基于模型的评估提供了一条有前景的路径。

英文摘要

Multimodal large language models (MLLMs) are increasingly used as evaluators, yet their reliability in professional assessment tasks that require expert judgment remains unclear. We investigate this challenge in the context of assessing psychotherapeutic interpersonal skills and introduce MIS-Bench, a Multimodal Interpersonal Skills (MIS) benchmark comprising 996 psychotherapy response videos annotated across 8 dimensions of Facilitative Interpersonal Skills. Across 9 MLLMs with multiple modality and prompting settings, we find that current models show only modest agreement with human experts, inconsistent gains from multimodal input, and limited benefits from reasoning-based prompting. To mitigate this gap, we propose MIS-RAFT, a regression-aware fine-tuning method inspired by RAFT and tailored to fine-grained interpersonal skill scoring at one-decimal precision. MIS-RAFT addresses the mismatch between autoregressive token prediction and scalar-valued expert assessment, significantly improving agreement with human ratings. Overall, MIS-Bench reveals a clear gap between general multimodal capability and expert-level interpersonal judgment, while MIS-RAFT offers a promising path toward more reliable model-based assessment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑