HiMA-MDD:用于临床访谈中可解释多模态抑郁检测的分层多智能体框架
HiMA-MDD: A Hierarchical Multi-Agent Harness for Interpretable Multimodal Depression Detection in Clinical Interviews
浏览论文内容
中文总结 AI 辅助
本研究针对临床访谈多模态抑郁检测的分层评估需求,提出分层多智能体框架HiMA-MDD,其分三层智能体处理流程,在E-DAIC数据集上以Qwen2.5-72B-Instruct为骨干,性能优于现有最优方法。
中文摘要 AI 辅助
从多模态临床访谈中评估抑郁需整合来自多种症状的分散证据,形成连贯的PHQ-8量表概况。该过程具有分层性:相关证据在局部问答交互中往往稀疏且依赖上下文,多个交互共同支撑症状级判断,最终评估取决于完整症状概况的一致性。现有大语言模型(LLM)系统要么整体处理访谈,要么将工作分配给通用智能体角色,两种设计均未提供明确的协调机制,以在各层级间协调证据访问、条目评分权限、有限反馈及状态记录。为解决此差距,我们提出HiMA-MDD,一种分层多智能体框架,将该评估层级与三个智能体层对齐。非智能体预处理构建保留上下文的多模态问答单元后,第1层识别候选问答-条目关系并支持基于条目的有限证据路由;第2层将症状组分配给操作因子专家,每位专家负责一项临时条目评分;第3层审核完整临时概况,请求最多一轮针对性修订,并重建经验证的PHQ-8概况。该分层设计自然生成分层证据轨迹,保留所有中间证据、判断及修订以实现可审计性。最终条目评分确定性地生成总分及筛查决策。以Qwen2.5-72B-Instruct为框架主干,我们在E-DAIC数据集上的实验表明,HiMA-MDD的性能优于对比的现有最优方法。
英文摘要
Depression assessment from multimodal clinical interviews requires integrating dispersed evidence from multiple symptoms into a coherent PHQ-8 profile. This process is hierarchical: relevant evidence is often sparse and context-dependent within local question-answer exchanges, multiple exchanges jointly support symptom-level judgments, and the final assessment depends on the coherence of the complete symptom profile. Existing LLM systems either process interviews holistically or distribute work across generic agent roles; neither design necessarily provides an explicit orchestration mechanism that coordinates evidence access, item-score authority, bounded feedback, and state recording across these levels. To address this gap, we introduce HiMA-MDD, a hierarchical multi-agent harness that aligns this assessment hierarchy with three agent layers. After non-agentic preprocessing constructs context-preserving multimodal QA units, Layer 1 identifies candidate QA-to-item relations and supports bounded item-grounded evidence routing. Layer 2 assigns symptom groups to operational factor specialists, with one specialist responsible for each provisional item score. Layer 3 audits the complete provisional profile, requests at most one round of targeted revision, and reconstructs the verified PHQ-8 profile. This layered design naturally yields a Hierarchical Evidence Trace, preserves all intermediate evidence, judgments, and revisions for auditability. The final item scores then deterministically produce the total score and screening decision. Using Qwen2.5-72B-Instruct as the harness backbone, our experiments on E-DAIC demonstrate that HiMA-MDD outperforms the compared state-of-the-art methods.