arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04999cs.CL

BIT.UA 团队参与 BioASQ 14B:基于 pg_textsearch 和 Qdrant 的模块化检索及智能体驱动的答案生成

BIT.UA at BioASQ 14B: Modular Retrieval with pg_textsearch and Qdrant, and Agent-Based Answer Generation

André Ribeiro, Rúben Garrido, Alexander Christiansen, Richard A. A. Jonker, Sérgio Matos

首次发表
浏览论文内容

中文总结 AI 辅助

BIT.UA 团队参与 BioASQ 14B 挑战赛,通过模块化检索与智能体表决机制优化流水线,在多阶段任务中取得竞争力结果,还首次参与片段生成子任务。

中文摘要 AI 辅助

本文介绍了来自阿威罗大学的 BIT.UA 团队参与第 14 届 BioASQ Task B 生物医学问答挑战赛的情况。在以往提交成果的基础上,我们引入了大幅重构后的模块化代码库,并对流水线的检索和生成组件进行了重大改进。对于 A 阶段文档检索,我们将 PyTerrier PISA 索引替换为基于 PostgreSQL 的 pg_textsearch 以实现 BM25 检索,并采用 Qdrant 进行稠密嵌入索引,从而实现更高效的存储和 GPU 加速的相似度搜索。我们探索了基于 HyDE 的查询扩展策略以及 Context-1 检索方法。我们开发了新的重排器训练流水线,其中包含用于负采样的稠密检索。对于 A+ 阶段和 B 阶段的答案生成,我们引入了 LLM-as-a-judge 框架和一种新型智能体表决机制,即多个具有不同提示的智能体通过辩论并利用自适应文档保留机制迭代收敛至共识答案。我们还首次参与了片段生成子任务。我们的系统在所有批次中均取得了具有竞争力的结果,其中 A 阶段系统在第 1、3 批次中达到了 MAP 排名 5。我们讨论了这些架构变更的影响、经验教训,并概述了未来工作方向,包括集成 SPLADE 和 ColBERT。所有代码均公开可用。

英文摘要

This paper describes the participation of the BIT.UA team from the University of Aveiro in the 14th edition of the BioASQ Task B challenge on biomedical question answering. Building on our previous submissions, we introduced a substantially refactored and modular codebase, and made significant changes to both the retrieval and generation components of the pipeline. For Phase~A document retrieval, we replaced the PyTerrier PISA index with PostgreSQL-based pg\_textsearch for BM25 retrieval and adopted Qdrant for dense embedding indexing, enabling more efficient storage and GPU-accelerated similarity search. We explored HyDE-based query expansion alongside a Context-1 retrieval strategy. A new reranker training pipeline was developed, incorporating dense retrieval for negative sampling. For Phases A+ and B answer generation, we introduced an LLM-as-a-judge framework and a novel agent quorum mechanism, where multiple agents with diverse prompts debate and iteratively converge on a consensus answer using adaptive document retention. We also participated in the snippets generation subtask for the first time. Our systems achieved competitive results across all batches, with Phase~A systems achieving MAP ranks of 5 (Batch~1,3). We discuss the impact of these architectural changes, lessons learned, and outline directions for future work including SPLADE and ColBERT integration. All code is openly available: https://github.com/bioinformatics-ua/BioASQ14b.

补充信息

↑