arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23435cs.LGcs.CL

工具增强的在线策略蒸馏用于序列组学任务中的大语言模型领域适配

Tool-Augmented On-Policy Distillation for LLM Domain Adaptation in Sequence-Based Omics Tasks

  • Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
  • Yazhouwan National Laboratory(崖州湾国家实验室)
  • Shanghai Innovation Institute(上海创新研究院)
  • The Chinese University of Hong Kong(香港中文大学)
  • Westlake University(西湖大学)

机构由 AI 辅助整理,请以论文原文为准。

Jie Ying, Zhefan Wang, Zihong Chen, Zhengqing Li, Jinzhe Li, Gang Li, Jian Liu, Fang Hu, Tao Luo, Zhonghang Yuan, Wanli Ouyang, Stan Z. Li, Fan Yang, Nanqing Dong

AI总结:

针对多组学序列任务,提出首个推理基准OmicsBench和工具增强的在线策略蒸馏方法TA-OPD,以增强大语言模型基于证据的生物学推理能力,提升预测性能。

AI中文摘要:

多组学序列包含复杂的生物模式,然而解读其机制以实现自动化科学发现仍然具有挑战性。随着大语言模型(LLMs)解读这些序列,评估预测和科学推理变得至关重要。然而,现有的多组学序列任务基准依赖于分类和回归指标,忽略了模型是否掌握潜在生物学证据。我们引入了OmicsBench,这是首个针对多组学序列的推理基准,包含1,160个由专家验证的问题,涵盖DNA调控、RNA加工和蛋白质功能等六项任务。OmicsBench要求可追溯的证据链,并使用与领域专家共同开发的实例特定评分标准进行评估。对17个LLMs的评估揭示了一个反向关系:虽然科学LLMs在序列分类准确性上优于通用LLMs,但它们未能提供有效证据来支持其预测。一个合理的解释是捷径学习:专门模型可能依赖统计模式而非科学发现所需的生物机制。受此发现启发,我们引入了工具增强的在线策略蒸馏(TA-OPD),一种后训练方法,将序列预测与基于证据的生物学推理对齐。在跨越0.8B到27B参数的五个Qwen3.5模型上,TA-OPD持续增强生物学证据基础,同时在大多数任务上提高预测性能。这些收益在不同模型规模中持续存在,表明更强的序列推理并非仅来自模型容量的增加,而是可以通过证据感知训练得到改进。总之,OmicsBench和TA-OPD提供了一个诊断多组学LLMs推理失败的框架,以及一条通往预测更好地基于生物学意义证据的模型的路径。

英文摘要:

Multi-omics sequences contain complex biological patterns, yet deciphering their mechanisms for automated scientific discovery remains challenging. As large language models (LLMs) interpret these sequences, evaluating both predictions and scientific reasoning is critical. However, existing benchmarks for multi-omics sequence tasks rely on classification and regression metrics, neglecting whether models grasp the underlying biological evidence. We introduce OmicsBench, the first reasoning benchmark for multi-omics sequences, comprising 1,160 expert-validated questions across six tasks spanning DNA regulation, RNA processing, and protein function. OmicsBench requires traceable evidence chains, evaluated using instance-specific rubrics developed with domain experts. Evaluating 17 LLMs reveals an inverse relationship: while scientific LLMs outperform general-purpose LLMs in sequence classification accuracy, they fail to provide valid evidence to support their predictions. One plausible interpretation is shortcut learning: specialized models may rely on statistical patterns rather than the biological mechanisms needed for scientific discovery. Motivated by this finding, we introduce tool-augmented on-policy distillation (TA-OPD), a post-training method to align sequence prediction with evidence-grounded biological reasoning. Across five Qwen3.5 models spanning 0.8B to 27B parameters, TA-OPD consistently strengthens biological evidence grounding while improving predictive performance on most tasks. These gains persist across model scales, indicating that stronger sequence reasoning does not arise solely from increased model capacity, but can be improved through evidence-aware training. Together, OmicsBench and TA-OPD provide a framework for diagnosing reasoning failures in multi-omics LLMs and a path toward models whose predictions are better grounded in biologically meaningful evidence.

补充信息

↑