DepressionAgent:用于抑郁风险评估的阅读、倾听、观察与审议多模态证据框架
DepressionAgent: Reading, Listening, Seeing, and Deliberating Multimodal Evidence for Depression Risk Assessment
浏览论文内容
中文总结 AI 辅助
研究针对现有多模态抑郁评估隐式特征融合的不足,提出以证据为中心的DepressionAgent框架,通过显式证据审议等机制在多基准上取得竞争力性能,且有效性与可检查性获多维度验证。
中文摘要 AI 辅助
多模态抑郁风险评估需要联合解读文本、声学和视觉线索,这些线索往往细微、非特定、依赖上下文,且跨模态间可能存在不一致。现有多模态方法主要通过特征融合学习隐式表示,使得预测背后的证据及跨模态分歧的处理在很大程度上是隐式的。我们提出DepressionAgent,一个以证据为中心的智能体框架,将多模态抑郁评估从隐式特征融合转变为显式证据审议。DepressionAgent首先将文本、声学和视觉输入转换为特定模态的证据,然后将自我报告和行为证据组织成并行的支持-挑战审议分支。跨模态仲裁明确检查两个分支之间的一致性与分歧,冲突反思在决策前重新审视不一致的评估。后续的风险反思机制为初始低风险案例提供独立的文本二次意见,以减少潜在的漏检风险信号。在未进行抑郁特定监督训练或参数微调的情况下,DepressionAgent在多个公共基准上取得了有竞争力的性能。大量的 ablation( ablation: ablation实验,即消融实验)、跨模型评估、定性分析及临床医生评估进一步证明了所提框架的有效性与可检查性。
英文摘要
Multimodal depression risk assessment requires jointly interpreting textual, acoustic, and visual cues that are often subtle, non-specific, context-dependent, and potentially inconsistent across modalities. Existing multimodal approaches predominantly learn latent representations through feature fusion, leaving the evidence underlying a prediction and the treatment of cross-modal disagreement largely implicit. We propose DepressionAgent, an evidence-centric agentic framework that transforms multimodal depression assessment from implicit feature fusion into explicit evidence deliberation. DepressionAgent first converts textual, acoustic, and visual inputs into modality-specific evidence, and then organizes self-report and behavioral evidence into parallel support--challenge deliberation branches. Cross-modal arbitration explicitly examines agreement and disagreement between the two branches, with conflict reflection revisiting inconsistent assessments before decision making. A subsequent risk reflection mechanism provides an independent textual second opinion for initially low-risk cases to reduce potentially missed risk signals. Without depression-specific supervised training or parameter fine-tuning, DepressionAgent achieves competitive performance on multiple public benchmarks. Extensive ablations, cross-model evaluations, qualitative analyses, and clinician assessments further demonstrate the effectiveness and inspectability of the proposed framework.