arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28557cs.AIq-bio.GN

BaseCamp——一个用于自动化DNA测序数据管道的智能体AI框架

BaseCamp --- An Agentic AI Framework for Automating DNA Sequencing Data Pipelines

Eranga Bandara, Xueping Liang, Asanga Gunaratna, Tharaka Hewa, Abdul Rahman, Peter Foytik, Safdar H. Bouk, Sachini Rajapakse, Isurunima Kularathna, Pramoda Karu… 展开作者

Eranga Bandara, Xueping Liang, Asanga Gunaratna, Tharaka Hewa, Abdul Rahman, Peter Foytik, Safdar H. Bouk, Sachini Rajapakse, Isurunima Kularathna, Pramoda Karunarathna, Chalani Rajapakse, Ng Wee Keong, Kasun De Zoysa, Amin Hass, Wathsala Herath, Ross Gore, Ravi Mukkamala, Nihal Siriwardanagea, Gihan Siriwardanagea, Aruna Withanage, Nilaan Loganathan, Sachin Shetty

首次发表
浏览论文内容

中文总结 AI 辅助

BaseCamp 提出一个智能体 AI 框架,将 DNA 测序管道的决策层分解为六个专门智能体,协调本地微调 LLM 自动配置工具、解释结果并检测异常,在保持可重复性的同时实现专家级自动化。

中文摘要 AI 辅助

DNA测序管道涵盖质量控制、比对、变异 calling 和注释,现已由工作流管理系统可靠执行,这些系统能够大规模编排已建立的生物信息学工具。仍然需要人工操作的是围绕该执行的决策层:选择适合样本和平台的质量阈值、裁定边缘变异 calling、诊断异常,以及确定哪些发现值得专家审查。这些决策重复性强、判断密集、在不同操作者之间不一致,且经常缺乏文档记录。本文介绍了 BaseCamp,一个用于自动化 DNA 测序管道决策层的新型智能体 AI 框架。该框架将管道分解为六个专门的 AI 智能体,涵盖样本接收与质量控制、比对、变异 calling、注释、跨阶段监控和报告。关键的是,BaseCamp 智能体不执行序列分析:已建立的工具执行比对、calling 和注释,而智能体在这些工具中进行选择、配置它们、解释其输出,并决定后续步骤。这将语言模型推理限制在判断层,在该层中推理可靠,并保留了现有工具所保证的可重复性。智能体推理由一组经过微调、领域专门化的大型语言模型组成的联盟驱动,并由一个中央推理 LLM 协调,在本地执行,因此测序数据不会离开操作环境,并在人机回环编排下进行。评估显示,智能体生成的配置与专家实践一致,显式的过滤账本使过滤过程可检查,而跨阶段异常检测能够发现执行监控遗漏的情况。BaseCamp 为科学数据管道的智能体自动化提供了可推广的蓝图。

英文摘要

DNA sequencing pipelines, spanning quality control, alignment, variant calling, and annotation, are now reliably executed by workflow management systems that orchestrate established bioinformatics tools at scale. What remains manual is the decision layer surrounding that execution: selecting quality thresholds appropriate to a sample and platform, adjudicating borderline variant calls, diagnosing anomalies, and determining which findings warrant expert review. These decisions are repetitive, judgment-intensive, inconsistent across operators, and frequently undocumented. This paper introduces BaseCamp, a novel agentic AI framework for automating the decision layer of DNA sequencing pipelines. The framework decomposes the pipeline into six specialized AI agents, covering sample intake and quality control, alignment, variant calling, annotation, cross-stage monitoring, and reporting. Critically, BaseCamp agents do not perform sequence analysis: established tools execute alignment, calling, and annotation, while the agents select among them, configure them, interpret their output, and decide what follows. This confines language model reasoning to the judgment layer where it is reliable and preserves the reproducibility existing tooling guarantees. Agent reasoning is powered by a consortium of fine-tuned, domain-specialized large language models coordinated by a central reasoning LLM, executing locally so no sequencing data leaves the operating environment, under human-in-the-loop orchestration. Evaluation shows agent-generated configurations are concordant with expert practice, that an explicit filtering ledger renders inspectable what filtering otherwise removes without trace, and that cross-stage anomaly detection surfaces conditions execution monitoring misses. BaseCamp offers a generalizable blueprint for agentic automation of scientific data pipelines.

发表机构

  • Old Dominion University(老道明大学)
  • Deloitte & Touche LLP(德勤会计师事务所)
  • AI Motion Labs(AI Motion 实验室)
  • University of Oulu(奥卢大学)
  • Nanyang Technological University(南洋理工大学)
  • University of Colombo(科伦坡大学)
  • Accenture Technology Labs(埃森哲技术实验室)
  • GSI Scandinavia AB(GSI 斯堪的纳维亚公司)

机构由 AI 辅助整理,请以论文原文为准。

↑