SSE-Bio:面向多跳生物医学推理的、具备智能体式检索策略的结构化自进化智能体
SSE-Bio: A Structured Self-Evolving Agent with Agentic Retrieval Policy for Multi-Hop Biomedical Reasoning
- School of Computing Science, University of Glasgow(格拉斯哥大学计算科学学院)
- Brigham and Women’s Hospital, Harvard Medical School, Harvard University(哈佛大学医学院附属布莱根妇女医院)
- Language Technology Lab, University of Cambridge(剑桥大学语言技术实验室)
- School of Natural and Computing Science, University of Aberdeen(阿伯丁大学自然与计算科学学院)
- School of Cancer Sciences, University of Glasgow(格拉斯哥大学癌症科学学院)
- Cancer Research UK Scotland Institute(英国癌症研究苏格兰研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对生物医学多跳问答的指令漂移问题,提出SSE-Bio结构化自进化智能体,通过代理策略选择性检索、细粒度模板编辑及组相对策略优化,在三个基准上实现优于现有基线的性能,BioHopR上提升6.56绝对点。
AI中文摘要:
生物医学多跳问答(QA)要求模型连接疾病、药物、蛋白质、表型等中间实体的证据。现有智能体通常依赖静态检索工作流或粗粒度提示重写,当推理流程需更新时会导致指令漂移。我们提出SSE-Bio,一种具备智能体式检索策略的结构化自进化智能体,用于多跳生物医学推理。SSE-Bio不全局重写智能体指令,而是维护结构化状态,通过可训练的代理策略选择性检索知识三元组与先验模板,并通过细粒度模板编辑提升推理记忆。为优化检索决策,我们引入基于组相对策略优化的代理训练策略,其中代理通过替代检索选择的决策对比组进行改进。在三个生物医学多跳QA基准上的实验表明,SSE-Bio始终优于现有基线,在BioHopR上比最强的自进化基线实现了6.56个绝对点的提升。
英文摘要:
Biomedical multi-hop question answering (QA) requires models to connect evidence across intermediate entities such as diseases, drugs, proteins, and phenotypes. Existing agents typically rely on static retrieval workflows or coarse-grained prompt rewriting, which can lead to instruction drift when reasoning procedures need to be updated. We propose SSE-Bio, a structured self-evolving agent with an agentic retrieval policy for multi-hop biomedical reasoning. Instead of globally rewriting agent instructions, SSE-Bio maintains a structured state, selectively retrieves knowledge triplets and prior templates through a trainable proxy policy, and improves its reasoning memory through fine-grained template editing. To optimise retrieval decisions, we introduce a proxy-training strategy based on group relative policy optimization, where the proxy is improved through decision-contrastive groups over alternative retrieval choices. Experiments on three biomedical multi-hop QA benchmarks show that SSE-Bio consistently outperforms existing baselines, achieving an improvement of 6.56 absolute points over the strongest self-evolving baseline on BioHopR.