发表机构
Stony Brook University; Renaissance School of Medicine at Stony Brook University(石溪大学; 石溪大学雷内桑斯医学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出ARGUS框架,将确定性生物学计算与LLM推理分离,针对疾病相关非编码区单核苷酸变异,通过多数据库证据分析,为不同转录因子生成差异化研究轨迹,解决大语言模型解释变异时的幻觉问题。
AI 中文摘要
全基因组关联研究中超过90%的疾病相关变异位于非编码调控区域,但其功能解释仍是基因组医学领域的核心开放性问题。用于解释此类变异的大语言模型常出现转录因子(TF)结合变化的幻觉、编造实验支持、对统计可忽略信号赋予生物学意义等问题。本文提出ARGUS(面向不确定性感知科学家的智能体调控基因组学),其严格将确定性生物学计算与大语言模型(LLM)介导的推理分离。ARGUS将458个基于DNABERT的TF结合模型封装在一个假设导向的研究循环中,其中规划器根据当前不确定性选择证据来源,验证器确定性地解释每个观测结果,中间结果会改变研究路径。针对8q24癌症风险位点的变异rs6983267,同一规划器为4个TF生成了4条不同轨迹:FOXA1在3步内被修正,因为真实的ADASTRA等位基因特异性结合数据(15个实验,FDR=0.030)揭示了被饱和掩盖的模型假阴性;KLF6在ADASTRA、JASPAR基序分析和ENCODE cCRE调控注释中遍历8步后,因混合间接证据弃权(不执行);RAD21在8步后弃权,因ADASTRA返回了覆盖度合格但无统计学意义的等位基因检验(5个实验,FDR=0.65);SP1与FOXA1共享饱和保留预测,因该位点无直接实验证据而弃权。所有观测结果均来自真实的ADASTRA、JASPAR和ENCODE cCRE查询,无模拟数据。对固定优先级规划与LLM介导规划的比较显示,LLM规划器通过拒绝无法解决所检验主张的证据,以更少的工具调用得出了相同结论。
英文摘要
Over 90% of disease-associated variants from genome-wide association studies fall in noncoding regulatory regions, yet their functional interpretation remains a central open problem in genomic medicine. Large language models prompted to interpret such variants routinely hallucinate transcription factor (TF) binding changes, fabricate experimental support, and assign biological significance to statistically negligible signals. We present ARGUS (Agentic Regulatory Genomics for an Uncertainty-aware Scientist), which strictly separates deterministic biological computation from LLM-mediated reasoning. ARGUS wraps 458 DNABERT-based TF binding models in a hypothesis-directed investigation loop where a planner selects evidence sources based on current uncertainty, a verifier deterministically interprets each observation, and intermediate results change the investigation path. On variant rs6983267 at the 8q24 cancer risk locus, the same planner produces four divergent trajectories for four TFs. FOXA1 is rescued in 3 steps when real ADASTRA allele-specific binding data (15 experiments, FDR = 0.030) reveals a model false negative masked by saturation. KLF6 traverses 8 steps across ADASTRA, JASPAR motif analysis, and ENCODE cCRE regulatory annotation before abstaining due to mixed indirect evidence. RAD21 abstains in 8 steps after ADASTRA returns a coverage-qualified but nonsignificant allelic test (5 experiments, FDR = 0.65), and SP1, which shares FOXA1's saturated retained prediction, abstains because no direct experimental evidence exists at this locus. All observations come from real ADASTRA, JASPAR, and ENCODE cCRE queries; none are simulated. A comparison of fixed-priority and LLM-mediated planning shows that the LLM planner reaches identical verdicts with fewer tool calls by declining evidence that cannot resolve the claim under test.
CommentsAccepted at the NeurIPS 2026 Workshop on Agentic AI for Biological Discovery (AgenticLS). Code: https://github.com/duttaprat/ARGUS