发表机构
Praxa Labs(普拉克斯实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究开发柏拉图生物,扩展开放架构并结合多种功能,修复评估缺陷。通过Python套件等验证其有效性,评估两个用例,如历史再发现任务及蛋白质结构比较,提供可重复软件契约和筛选基准,为生物研究提供支持。
AI 中文摘要
大型语言模型研究代理可以连接文献检索、分析代码和稿件准备,但连贯输出并不能确立科学有效性。我们开发了柏拉图生物,这是开放的柏拉图/德纳里奥架构的生物学导向扩展,它将明确的工作流程状态与出处记录、引用检查、主张与证据链接、限定范围的文件写入和出版门控相结合。一次源审核识别并修复了三个可能扭曲评估的缺陷。在当前的清洁版本中,完整的Python套件以931次通过、6次跳过完成,没有失败或错误;有针对性的生物学、基因组学、证据/引用和对抗安全套件同样顺利完成。我们评估了两个狭义用例。在一个冻结的历史再发现任务中,1986年之前的独立文献桥梁将鱼油与雷诺现象之间后来研究的关系排在首位;TF-IDF排在第二位,语料库频率排在第三位。在对15种人类蛋白质的AlphaFold模型与实验结构的单独比较中,11个靶点的高置信度核心C-αRMSD低于1埃(中位数为0.501埃)。四个靶点超过2埃,置信度掩码将SUMO1在74个残基上的差异从16.61降低到2.58埃。工作流程发出了27个可追溯的差异区域,所有这些都作为未经验证的假设保留下来。因此,柏拉图生物提供了可重复的软件契约和可审计的筛选基准;关于代理功效或生物新奇性的更广泛主张需要预先注册的评估、独立审查和前瞻性验证。
英文摘要
Large language model research agents can connect literature retrieval, analysis code, and manuscript preparation, but coherent output does not establish scientific validity. We developed Plato-Bio, a biology-routed extension of the open Plato/Denario architecture that couples explicit workflow states with provenance records, citation checks, claim-to-evidence links, scoped file writes, and publication gates. A source audit identified and repaired three defects that could distort evaluation: loss of task domain in the default factory, omission of declared method signals from scoring, and evidence sidecars that lacked the drafted-claim denominator. On the current clean revision, the full Python suite completed with 931 passes, six skips, and no failures or errors; targeted biology, genomics, evidence/citation, and adversarial-safety suites likewise completed without failure. We evaluated two narrow use cases. In a frozen historical rediscovery task, independent pre-1986 literature bridges ranked the later-studied relation between fish oil and Raynaud phenomenon first; TF-IDF ranked it second and corpus frequency third. This single curated task measures retrospective ranking, not prospective discovery. In a separate comparison of AlphaFold models with experimental structures for 15 human proteins, 11 targets had high-confidence-core C-alpha RMSD below 1 Angstrom (median 0.501 Angstrom). Four targets exceeded 2 Angstrom, and confidence masking reduced the SUMO1 discrepancy from 16.61 to 2.58 Angstrom over 74 residues. The workflow emitted 27 traceable discrepancy regions, all retained as unvalidated hypotheses. Plato-Bio therefore provides reproducible software contracts and auditable screening baselines; broader claims of agent efficacy or biological novelty require preregistered evaluation, independent review, and prospective validation.
Comments16 pages, 6 figures, 3 tables. Companion code and data: https://github.com/Eldergenix/Plato-Scientific-Research-Autonomous-Agent. This fork-specific validation study cites, but does not duplicate, arXiv:2510.26887