PEN-STACK:用于基因组写作的语言模型智能体的非制造工具层
PEN-STACK: A non-fabricating tool layer for language-model agents in genome writing
浏览论文内容
中文总结 AI 辅助
本研究提出PEN-STACK工具层,为基因组写作的语言模型智能体提供可靠来源的工具,可消除数量伪造,经测试能提升生物安全,为智能体基因组工程系统提供了开源基底。
中文摘要 AI 辅助
背景:语言模型智能体广泛应用于生物学领域,但它们报告的数量缺乏可验证来源,且存在未受管控的生物安全风险。基因组写作对上述两方面要求更为严格:编写计划必须明确位置、写入酶、货物及递送载体,所有要素均为定量且相互关联,若无集成工具层,智能体需自行提供这些信息。本文介绍PEN-STACK,这一开源工具层可提供带有可靠来源的上述信息。结果:PEN-STACK提供10个基因组写作设计阶段对应的22个范围感知工具,可通过软件开发工具包、模型上下文协议服务器及表述性状态传递接口访问,遵循类型强制不变量:每个数量必须源自经过验证的工具。无工具时,3种模型家族在朴素提示下伪造了240个所需数量中的90.8%至98.8%;经引导后,剩余伪造数量为0至4个,无模型达到零伪造认证。借助工具时,相同模型在四项目标审计中未伪造任何数量。排放前生物安全筛查对全部8个设计均匹配专家标签。表达-鲁棒性轴在精确位点分辨率下验证(ρ=0.571,n=1506),但在默认的较粗分辨率下未验证(ρ≈0.16),返回机器可读降级标志。10项预注册主张中有8项未通过,每项均被标记为机器可读。结论:证据表明,是接地而非提示或模型规模消除了伪造,且接地需要一个基底;接地的四项目标审计部分,需在240个字段规模上重复验证。PEN-STACK作为可导入的开源代码,为智能体基因组工程系统提供了该基底。
英文摘要
Background. Language-model agents are widely used in biology, but they report quantities without a verifiable source and pose unmanaged biosecurity risks. Genome writing sharpens both: a write plan must specify a location, writer enzyme, cargo, and delivery vehicle, all quantitative and interdependent, so without an integrated tool layer, the agent must supply them. We introduce PEN-STACK, an open tool layer that supplies them with guaranteed provenance. Results. PEN-STACK provides ten genome-writing design stages as twenty-two scope-aware tools, accessible via a software development kit, a Model Context Protocol server, and a Representational State Transfer interface, under a type-enforced invariant: every quantity must originate from a validated tool. Without tools, three model families fabricated 90.8% to 98.8% of the 240 required quantities under a naive prompt; coaching left a residual of 0 to 4, with no model certified at zero. Driving the tools, the same models fabricated nothing on a four-goal audit. A pre-emission biosecurity screen matched expert labels on all eight designs. The expression-robustness axis validated at exact-site resolution (ρ= 0.571, n = 1,506) but not at the coarser resolution served by default (ρ\approx 0.16), which returns a machine-readable downgrade flag. Eight of ten pre-registered claims did not pass, each flagged as machine-readable. Conclusions. On this evidence, grounding, not prompting or model scale, removes fabrication, and grounding requires a substrate; the grounded arm, a four-goal audit, warrants replication at the 240-field scale. PEN-STACK provides that substrate as open, importable code for agentic genome-engineering systems.