AI 中文总结
GenomeHarness是一种智能体式工具,通过受控搜索微调方案适配基因组语言模型,在52个模型-任务设置中多数提升了测试MCC,使下游适配更可靠可审计。
AI 中文摘要
预训练基因组语言模型为DNA序列分析提供了可复用的表征,但将其转化为可靠的下游预测器仍非易事。它们的实际性能高度依赖微调方案,而现有研究报告的默认方案可能对新任务或模型主干并非最优,导致下游结果薄弱且难以解释。这些要求给许多目标用户带来了巨大的操作负担,这些用户的专业知识通常集中在生物学问题和解释上,而非机器学习工程。因此,可靠使用基因组语言模型不仅需要传统的AutoML风格调优,还需要一种系统、预算感知且可审计的流程,以降低下游适配的门槛。我们提出了GenomeHarness,这是一种通过对微调方案进行受控搜索来适配基因组语言模型的智能体式工具。GenomeHarness结合了用于提出和修复方案编辑的AI智能体、用于协议约束执行、资源管理和测试隔离的工具,以及用于在方案谱系间分配搜索精力的蒙特卡洛树搜索控制器。我们在DNABERT2和NTv2-100M-Multi上,针对NT基准和基因组基准评估了GenomeHarness。在方案冻结后,使用3个随机种子进行最终评估。在52个模型-任务设置中,GenomeHarness在47个设置中提高了平均测试马修斯相关系数(MCC),包括26个DNABERT2设置中的24个和26个NTv2-100M-Multi设置中的23个。在基因组基准以及根方案不稳定或匹配不佳的任务(如人类Ensembl OCR任务)上,提升尤为显著。搜索轨迹进一步表明,GenomeHarness逐步识别出更强的方案,将下游适配转化为受控且可审计的工作流程,而非手动调优过程。
英文摘要
Pretrained genome language models provide reusable representations for DNA sequence analysis, but turning them into reliable downstream predictors remains non-trivial. Their practical performance depends strongly on fine-tuning recipes, and default recipes reported in prior studies may be suboptimal for new tasks or model backbones, making weak downstream results difficult to interpret. These requirements place a substantial operational burden on many intended users, whose expertise is often centered on biological questions and interpretation rather than machine-learning engineering. Reliable use of genome language models therefore requires more than conventional AutoML-style tuning: it requires a systematic, budget-aware, and auditable procedure that lowers the barrier to downstream adaptation. We present GenomeHarness, an agentic harness for adapting genome language models through controlled search over fine-tuning recipes. GenomeHarness combines an AI agent for proposing and repairing recipe edits, a harness for protocol-constrained execution, resource management, and test isolation, and a Monte Carlo tree search controller for allocating search effort across recipe lineages. We evaluate GenomeHarness on DNABERT2 and NTv2-100M-Multi across the NT Benchmark and Genomic Benchmarks. Final evaluation is performed using three random seeds after recipe freezing. Across 52 model-task settings, GenomeHarness improves mean test MCC in 47 settings, including 24 of 26 DNABERT2 settings and 23 of 26 NTv2-100M-Multi settings. The gains are especially pronounced on Genomic Benchmarks and on tasks where the root recipe is unstable or poorly matched, such as human ocr ensembl task. Search traces further show that GenomeHarness progressively identifies stronger recipes, turning downstream adaptation into a controlled and auditable workflow rather than a manual tuning process.
Comments10 pages, 4 figures