发表机构
MIT CSAIL; Massachusetts General Hospital; Harvard Medical School(麻省理工学院计算机科学与人工智能实验室; 麻省总医院; 哈佛医学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文评估了多模态大语言模型在自动痴呆症分类中的推理能力,发现基于文本理由的策略会导致幻觉和不一致,并提出DeTAiL框架,通过非线性适配器和强化学习利用内部表示,在两个数据集上优于基线方法。
AI 中文摘要
多模态大语言模型(MLLMs)已成为一种有前景的方法,用于提高从语音录音中自动痴呆症分类(ADC)系统的准确性、可迁移性和可解释性。然而,它们的推理能力是否对ADC有益,以及如何利用这些能力仍不清楚。在本文中,我们对用于ADC的推理MLLMs进行了仔细评估,并表明依赖基于文本理由等朴素策略可能导致诊断理由的幻觉和不一致,并且与无LLM的基线相比,ADC性能较差。为了克服这一限制,我们提出了\textbf{De}mentia \textbf{T}hinker with Nonlinear \textbf{A}daptor and Re\textbf{i}nforcement \textbf{L}earning (DeTAiL),一种基于适配器的框架,利用推理MLLMs的内部表示来改进痴呆症分类。在两个具有不同测试格式和标签粒度的痴呆症数据集上,DeTAiL始终优于强基线和依赖基于文本理由的方法。代码和演示将在接收后发布。
英文摘要
Multimodal large language models (MLLMs) have emerged as a promising approach for improving the accuracy, transferability, and explainability of automatic dementia classification (ADC) systems from voice recordings. Yet it remains unclear whether their reasoning capabilities are beneficial for ADC, and how such capabilities should be leveraged. In this paper, we conduct a careful evaluation of reasoning MLLMs for ADC and show that naive strategies, such as relying on text-based rationales, can lead to hallucinated and inconsistent rationales for diagnosis and yield inferior ADC performance compared with LLM-free baselines. To overcome this limitation, we propose \textbf{De}mentia \textbf{T}hinker with Nonlinear \textbf{A}daptor and Re\textbf{i}nforcement \textbf{L}earning (DeTAiL), an adaptor-based framework that exploits the internal representations of reasoning MLLMs for improved dementia classification. Across two dementia datasets with distinct test formats and label granularities, DeTAiL consistently outperforms strong baselines and methods that rely on text-based rationales. Code and demo will be released upon acceptance.