发表机构
University of Florida; University of Chicago; Data Science Institute, University of Chicago; Toyota Technological Institute at Chicago; Mayo Clinic; Argonne National Laboratory(佛罗里达大学; 芝加哥大学; 芝加哥大学数据科学研究所; 芝加哥丰田技术研究所; 梅奥诊所; 阿贡国家实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对DNA/RNA等生物分子设计难题,提出MCTH框架,通过蒙特卡洛树搜索实现序列-结构协同设计,在多模态设计任务中性能优于基线,且可跨模态泛化。
AI 中文摘要
生物分子设计支撑着从分子识别到治疗药物及合成生物学的各类应用,但全新交互设计仍具挑战性——尤其对于DNA/RNA这类研究不足的非蛋白模态,其复杂数据稀缺且异质,几何与化学约束更严苛。我们提出MCTH(蒙特卡洛树幻觉,Monte Carlo Tree Hallucination),这是一种仅用于推理的框架,将全原子序列-结构协同设计转化为基于预训练折叠与逆折叠模型生成的幻觉状态、且具备不确定性感知的规划,可在同一决策循环中实现可选的生物物理控制。MCTH将这些模型视为冻结的黑箱算子,使用蒙特卡洛树搜索在竞争的设计轨迹间分配固定推理预算,纳入模型置信度与不确定性,以及存在多个预测器时的跨专家共识/分歧。在蛋白质-RNA、蛋白质-DNA、蛋白质-蛋白质及蛋白质-配体设计任务中,匹配预算的实验显示,自适应搜索优于更简单的采样与循环策略; hold-out的AlphaFold3与Chai-1评估表明,其可泛化至搜索时的先验之外。MCTH提供了跨模态的共享规划层,同时支持任务特定的折叠、逆折叠及生物物理模块,无需对组件模型进行微调或反向传播。
英文摘要
Biomolecular design underpins applications from molecular recognition to therapeutics and synthetic biology, yet de novo interaction design remains challenging-especially for DNA/RNA, underexplored non-protein modalities with scarce, heterogeneous complex data and sharper geometric and chemical constraints. We introduce MCTH (Monte Carlo Tree Hallucination), an inference-only framework that casts all-atom sequence-structure co-design as uncertainty-aware planning over hallucinated states from pretrained folding and inverse-folding models, with optional biophysical control within the same decision loop. MCTH treats these models as frozen black-box operators and uses Monte Carlo Tree Search to allocate a fixed inference budget across competing design trajectories, incorporating model confidence and uncertainty, as well as cross-expert consensus/disagreement when multiple predictors are available. Across protein-RNA, protein-DNA, protein-protein, and protein-ligand design, matched-budget experiments show that adaptive search improves over simpler sampling and cycling strategies, while held-out AlphaFold3 and Chai-1 evaluations demonstrate transfer beyond the search-time oracle. MCTH provides a shared planning layer across modalities while allowing task-specific folding, inverse-folding, and biophysical modules, requiring no fine-tuning or backpropagation through component models.