arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Bioinfoysis 技术报告

Bioinfoysis Technical Report

Qingyang Shao, Xin Zhang, Zhouyang Yuan, Xianying Chen, Yujia Xiang, Zihao Yang, Tong Ye, Yangqi Zhang, Jiakang Xu, Xiaoqing Yan, Xuan Luo, Keyi Li, Enci Fan, Kai Kang, Zhuohan Liu, Xingyu Jin, Chunran Teng, Tao Li, Xinyu Lyu, Minghui Wang, Wenfeng Li, Yidan Gao, Siyu Liu, Mingrui Luo, Zhu Liang, Guanren Qiao, Zhiping Xu

arXiv 2609.03871首次发表:更新:

发表机构

Bioinfoysis(Bioinfoysis)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出多智能体框架 Bioinfoysis,结合全局规划与证据驱动的逐步重规划,在 BixBench 等基准测试中大幅提升生物信息学问答任务准确率,为生物信息学自动化提供新方案。

AI 中文摘要

大型语言模型智能体在生物信息学领域展现出应用前景,但现有多数系统主要聚焦于生成最终答案,将规划、工具使用和代码执行视为短暂交互。这种设计不适用于长周期生物信息学任务,此类任务的结论必须与支撑它们的数据、计算及中间证据保持关联。我们推出 Bioinfoysis,这是一个多智能体框架,它将每个请求表示为基于持久工件的分析运行。Bioinfoysis 结合全局规划与逐步、证据驱动的重新规划:规划器维护可执行的检查清单,并在每个工作智能体执行后使用返回的结构化交接内容修订待处理步骤。这些交接内容将中间结果与其负责的智能体、检查清单步骤及规划生成绑定,防止重新规划后静默复用过时证据。受控运行时会在生成的脚本、表格和图形用于下游分析或报告前对其进行验证,而特定角色的上下文、持久记忆及受控生物信息学技能可支持在长分析轨迹上的可靠执行。我们在 BixBench 和 LAB-Bench 2 的两个问答任务赛道上对 Bioinfoysis 进行评估:在 BixBench 上,Bioinfoysis 达到 82.4% 的最优准确率;在四个基础语言模型上,Bioinfoysis 将 SeqQA2 的平均准确率从 27.81% 提升至 64.13%,将 DbQA2 的平均准确率从 3.13% 提升至 31.25%。这些结果表明,可靠的生物信息学自动化不仅取决于模型能力,还取决于管控规划、执行、记忆及证据流的框架。我们希望 Bioinfoysis 的出现能在生物信息学社区的发展中发挥驱动和引领作用,其演示网站可通过此 https URL 访问。

英文摘要

Large language model agents have shown promise in bioinformatics, but most existing systems focus primarily on producing final answers, treating planning, tool use, and code execution as transient interactions. This design is poorly suited to long-horizon bioinformatics tasks, where conclusions must remain connected to the data, computations, and intermediate evidence that support them. We introduce \textbf{Bioinfoysis}, a multi-agent harness that represents each request as a persistent, artifact-grounded analysis run. Bioinfoysis combines global planning with step-wise, evidence-driven replanning: the planner maintains an executable checklist and revises pending steps using structured handoffs returned after each worker execution. These handoffs bind intermediate results to their responsible agent, checklist step, and plan generation, preventing stale evidence from being silently reused after replanning. A controlled runtime validates generated scripts, tables, and figures before they are used in downstream analysis or reporting, while role-specific context, persistent memory, and governed bioinformatics skills support reliable execution over long analysis trajectories. We evaluate Bioinfoysis on BixBench and two question-answering tracks of LAB-Bench 2. On BixBench, Bioinfoysis achieves state-of-the-art accuracy of 82.4\%. Across four underlying language models, Bioinfoysis increases average accuracy from 27.81\% to 64.13\% on SeqQA2 and from 3.13\% to 31.25\% on DbQA2. These results demonstrate that reliable bioinformatics automation depends not only on model capability, but also on the harness that governs planning, execution, memory, and evidence flow. We hope that the emergence of Bioinfoysis will play a driving and leading role in the development of the bioinformatics community. Our demo website can be seen in https://report.bioinfoysis.com/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑