arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

佐治亚理工学院DS团队参加eRisk 2026任务3:针对ADHD症状句子的稀疏、语义及大语言模型重排序

DS@GT-ARC at eRisk 2026 Task 3: Sparse, Semantic, and LLM Reranking for ADHD Symptom Sentences

David Guecha

arXiv 2608.03883首次发表:更新:

AI 中文总结

本文介绍佐治亚理工学院DS团队针对eRisk 2026任务3,采用分阶段检索结合稀疏BM25、语义及LLM重排序的方案,其中LLM重排序器取得官方最优成绩,分阶段重排序具发展前景。

AI 中文摘要

本文介绍了我们参与eRisk 2026任务3(ADHD症状句子排序)的提交方案。该任务要求系统根据候选Reddit句子与成人ADHD自我报告量表(ASRS-v1.1)的18种症状的相关性对其进行排序。由于该任务首版未发布带标注的训练数据,我们依赖零样本实验、人工验证以及无监督或弱指导检索流程。我们的系统结合了稀疏BM25检索、针对自我指涉症状报告的证据感知重评分、基于嵌入的重排序、查询原型扩展以及基于大语言模型(LLM)的重排序。所有提交的系统均采用分阶段检索设计:BM25大规模检索候选,语义或LLM重排序器优化最终排序。在我们的提交方案中,LLM重排序器取得了最强的官方分数,其次是查询原型扩展方案。我们对排名前10的人工分析与官方专家评分趋势一致,表明分阶段重排序是值得进一步发展的有前景方向。

英文摘要

This paper describes our submissions to eRisk 2026 Task 3, ADHD Symptom Sentence Ranking. The task requires systems to rank candidate Reddit sentences according to their relevance to each of the 18 symptoms in the Adult ADHD Self-Report Scale (ASRS-v1.1). Because no annotated training data were released for this first edition of the task, we relied on zero-shot experimentation, manual validation, and unsupervised or weakly guided retrieval pipelines. Our systems combine sparse BM25 retrieval, evidence-aware rescoring for self-referential symptom reports, embedding-based reranking, query-prototype expansion, and LLM-based reranking. All submitted systems follow a staged retrieval design in which BM25 retrieves candidates at scale and semantic or LLM rerankers refine the final rankings. Among our submissions, the LLM reranker achieved the strongest official scores, followed by the prototype query-expansion run. Our manual top-10 analysis aligned with the official expert scoring trend, suggesting that staged reranking is a promising direction for further development.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑