arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迈向用语言控制生物:细胞、类器官和生物机器人的提示条件干预的离线学习

Toward Controlling Biology with Language:Offline Learning of Prompt-Conditioned Interventions for Cells, Organoids, and Biobots

Nam H. Le, Douglas Blackiston, Michael Levin, Josh Bongard

arXiv 2610.02247首次发表:更新:

发表机构

University of Vermont; Tufts University; Allen Discovery Center(佛蒙特大学; 塔夫茨大学; 艾伦发现中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究证明,利用视觉语言模型的判断作为唯一奖励,可离线学习将自然语言指令映射到生物干预的接口,并在异种机器人上实现80%的保留准确率。

AI 中文摘要

人工智能日益成为复杂技术系统的自然语言接口,让人们通过描述想要什么而非指定如何做来完成复杂的任务。将这种接口扩展到生命系统则更为困难:与代码或图像不同,生物干预没有封闭形式的语言含义,而学习这种映射所需的配对语言-干预-结果数据收集成本高昂,因为每个示例都需要其自身的湿实验室实验。一种绕过此问题的方法是将现有的干预及其已观察结果的档案视为固定的离线数据集,并使用视觉语言模型在不进行任何新实验的情况下判断存档结果是否与自然语言描述匹配。但是,这种判断是否足够可靠,足以在没有新实验且没有人工验证的情况下训练语言到干预的映射,此前仍不清楚。在此,我们表明,一个生命有机体——异种机器人(xenobot),一种没有神经系统的合成多细胞构造——的自然语言接口可以完全以这种方式离线学习,使用视觉语言模型自身的判断作为唯一的训练奖励:一条指令被映射到档案中已记录为产生所述行为的干预措施。这种映射泛化到全新的指令,并针对训练中保留的档案数据进行评估(80.0%的保留准确率,对比66.7%的随机基线,与直接基于真实标签训练的网络相匹配)。

英文摘要

Artificial intelligence increasingly serves as a natural-language interface to complex technical systems, letting people accomplish sophisticated tasks by describing what they want rather than specifying how to do it. Extending this interface to living systems is harder: unlike code or images, a biological intervention has no closed-form linguistic meaning, and the paired language-intervention-outcome data needed to learn such a mapping is expensive to collect, since each example requires its own wet-lab experiment. One way around this is to treat an existing archive of interventions and their already-observed outcomes as a fixed, offline dataset, and use a vision-language model to judge, without any new experiments, whether an archived outcome matches a natural-language description. But whether that judgment is reliable enough to train a language-to-intervention mapping on -- without new experiments and without human validation -- has remained unclear. Here we show that a natural-language interface for a living organism -- a xenobot, a synthetic multicellular construct with no nervous system -- can be learned entirely offline this way, using a vision-language model's own judgment as the sole training reward: an instruction is mapped to the intervention already on record as producing the described behavior. This mapping generalizes to entirely new instructions, evaluated against archive data withheld from training (80.0% held-out accuracy vs a $66.7% chance baseline, matching a network trained directly on ground-truth labels).

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑