DS@GT在eRisk 2026的ARC:用于对话式抑郁症筛查的具有结构化算法指导的混合多智能体大语言模型系统
DS@GT ARC at eRisk 2026: Hybrid Multi-Agent LLM System with Structured Algorithmic Guidance for Conversational Depression Screening
浏览论文内容
中文总结 AI 辅助
介绍DS@GT参加eRisk 2026对话式抑郁症筛查挑战赛,其系统历经三个阶段,最终采用混合配置并添加算法组件。提交三次运行结果,混合运行3表现出色,以低成本超付费基线,验证较弱开源模型经算法监督可与专有模型竞争。
中文摘要 AI 辅助
我们描述了DS@GT提交给eRisk 2026对话式抑郁症筛查任务1挑战赛的内容。在该挑战赛中,系统要访谈模拟不同抑郁症特征个体的大语言模型角色,并给出贝克抑郁量表II(BDI-II)得分及每个角色的四个关键症状,且不直接询问敏感心理健康问题。我们的流程历经三个阶段:从整体单模型原型开始,到基线多智能体架构,再到最终混合配置,用开源的Gemma 27B取代付费的GPT-5-nano访谈器。为弥补模型推理和指令遵循能力较弱的问题,混合配置添加了三个算法组件。我们提交了针对所有20个角色的三次全自动运行结果。混合运行3的ADODL为0.9063,在所有完整提交运行中排名第三,在21个团队中DS@GT整体排名第二,且以约四分之一的人均API成本超过付费基线运行1。这些结果支持了我们的核心假设,即通过足够的算法监督,较弱的开源模型能在对话访谈角色中与较强的专有模型竞争。我们的源代码可在该https网址获取。
英文摘要
We describe DS@GT's submission to the eRisk 2026 Task 1 challenge on conversational depression screening, in which systems interview LLM personas that simulate individuals with varying depression profiles and produce a Beck Depression Inventory II (BDI-II) score plus four key symptoms per persona, without directly asking sensitive mental health questions. Our pipeline evolved through three stages: a monolithic single-model prototype to start off, a baseline multi-agent architecture that separates conversational interviewing from BDI-II scoring under a coordinating orchestration layer, and a final hybrid configuration that replaces the paid GPT-5-nano interviewer with the open-source Gemma 27B. To offset the model's weaker reasoning and instruction-following, the hybrid adds three algorithmic components: a precomputed dialogue tree that standardizes interview openers and follow-ups, a reliability-weighted consensus aggregation inspired by the Weaver framework, and a cluster-based imputation step for unprobed symptoms. We submitted three fully automated runs across all 20 personas, with Run 1 from the paid baseline and Runs 2 and 3 from the hybrid. Hybrid Run 3 achieved an ADODL of 0.9063, ranking 3rd among all complete-submission runs and placing DS@GT 2nd among the 21 teams overall, while outperforming our paid baseline Run 1 (0.8841) at roughly one-quarter of the per-persona API cost. These results support our central hypothesis that with sufficient algorithmic supervision, a weaker open-source model can compete with a stronger proprietary model in the conversational interviewer role. Our source code is available at https://github.com/dsgt-arc/erisk-task1-2026.
发表机构
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。