arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13551cs.CL

用于BioASQ A+和B阶段的成本实用质量门控和选择融合多模型组合器

Cost-Pragmatic Quality Gating and Selection-Fusion Multi-Model Combiners for BioASQ Phases A+ and B

  • University of Technology Sydney(悉尼科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Dima Galat, Marian-Andrei Rizoiu

AI总结:

该研究针对BioASQ任务14B 2026系统,围绕检索弱时的重新检索及多模型答案组合展开。采用混合与代理驱动管道结合质量门控,成本实用策略提升效果。还分解多模型集成,解析器在召回方面表现出色,团队在初步排行榜上成绩优异。

AI中文摘要:

我们描述了我们的BioASQ任务14B 2026系统。这项工作集中在两个设计决策上:当第一阶段检索较弱时积极重新检索的程度,以及如何组合多个语言模型答案。检索结合了两个并行管道——一个混合第一阶段(密集BGE+BM25+RRF,在BioASQ-13b历史存档上达到R@200 = 99.3%)和一个代理驱动管道,该管道在PubMed、欧洲PMC和iCite上分解问题,并通过BGE交叉编码器质量门控标记弱支持问题以进行选择性重新检索。在2024年任务12B验证中,一种成本实用的重新检索策略在列表F1和列表精度上显著击败了技能严格的基线,重新检索成本降低了12%。在验证和测试13B(不同问题集)中保持提示和模型不变,在BioASQ发布的黄金输入池上列表F1绝对上升了+0.132,这与大量检索空间一致。对于B阶段回答,我们将多模型集成提升分解为一个由每个问题预言界定的选择组件和一个聚合器可以超越的融合组件。这种分解在任何实验之前预测,作为评判的大语言模型在选择主导的指标(是/否、多参考ROUGE)上获胜,但在融合友好指标的召回组件(事实类排名第一、列表召回)上结构上不足。在2025年任务13B中,我们的同义词联合解析器在每个方面都赢得了列表召回,而GPT-5.5单独保持了列表F1领先,因为解析器更广泛的项目集牺牲了精度。在2026年任务14B初步排行榜上,我们的团队在八个(阶段x批次)排行榜中的三个组合精确汇总上排名第一,赢得了四个单独问题类型单元格,并在B阶段b3理想情况下排名第一。

英文摘要:

We describe our BioASQ Task 14B 2026 system. The work centers on two design decisions: how aggressively to re-retrieve when first-stage retrieval is weak, and how to combine multiple language-model answers. Retrieval unions two parallel pipelines - a hybrid first stage (dense BGE + BM25 + RRF, reaching R@200 = 99.3% on the BioASQ-13b historical archive) and an agent-driven pipeline that decomposes the question over PubMed, Europe PMC, and iCite - with a BGE cross-encoder quality gate flagging weakly-supported questions for selective re-retrieval. On Task 12B 2024 validation, a cost-pragmatic re-retrieval policy beats a skill-strict baseline significantly on list F1 and list precision, at 12% lower re-retrieval cost. Holding prompt and model fixed across val and test 13B (different question sets), list F1 rises by +0.132 absolute on the BioASQ-released gold-input pool, consistent with substantial retrieval-side headroom. For Phase B answering we decompose multi-model ensemble lift into a selection component bounded by the per-question oracle and a fusion component that aggregators can exceed. The decomposition predicts before any experiment that LLM-as-judge wins on selection-dominated metrics (yes/no, multi-reference ROUGE) but is structurally insufficient on the recall component of fusion-friendly metrics (factoid rank-1, list recall). On Task 13B 2025 our synonym-union resolver wins list recall on every head, while GPT-5.5 solo retains the list-F1 lead because the resolver's wider item set costs precision. On the Task 14B 2026 preliminary leaderboard our team places first on the combined-exact aggregate on three of the eight (phase x batch) leaderboards, wins four individual question-type cells, and takes #1 on Phase B b3 ideal.

补充信息

↑