arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CASCADE:一种用于经患者数据验证的下游扰动预测的智能体调控网络框架

CASCADE: An Agentic Regulatory Network Framework for Patient-Data-Validated Downstream Perturbation Prediction

Jose A. Bird

arXiv 2608.05359首次发表:更新:

AI 中文总结

该研究提出CASCADE智能体调控网络框架,基于ARACNe网络预测基因扰动下游转录效应,通过TCGA、METABRIC数据验证其MYC基因预测方向的准确性,还评估LLM智能体映射自然语言请求至MCP工具调用的表现,发现其在模糊查询时存在错误选择扰动类型的问题。

AI 中文摘要

CASCADE是一种智能体框架,可基于预先计算的ARACNe调控网络预测基因扰动的下游转录效应,通过MCP工具实现。现有研究通过检查预测基因是否为已知癌症基因(成员身份)来验证此类工具;我们则采用焦点基因拷贝数扩增作为剂量依赖性替代指标,对应敲低的反向指标,结合真实TCGA患者肿瘤数据,测试预测的变化方向是否与实际相符。针对MYC基因,CASCADE预测的敲低靶标在三种癌症类型中与真实扩增/非扩增肿瘤表达表现出强一致性:BRCA为90.0%,COAD为72.0%,STAD为85.7%,所有p值均小于0.0013,远高于置换基线,通过PAM50亚型控制后仍成立,并在独立队列(METABRIC,87.2%)中得到重复。通过Fisher精确检验与 curated MSigDB基因集基线对比,CASCADE的准确率未显示超出现有MYC或E2F驱动生物学的公开知识,但它的基因特异性方向预测明显优于朴素均匀猜测。扩展至另外15个基因的验证显示,结果具有基因特异性而非普遍性:增殖机制调控因子大多可重复,而谱系身份转录因子及一个周期蛋白D旁系同源物(CCND2)始终失败,我们将此模式作为一种有保留的事后假设进行讨论。我们还单独评估了基于大语言模型(LLM)的智能体能否将自然语言请求正确映射为CASCADE的真实MCP工具调用。在35个查询中,已记录的本地模型达到71.4%的精确匹配率(更大模型为85.7%);模式与基因别名的错误可通过规模或服务器端修正解决,但两种模型在模糊查询时均会自信地默认选择错误的扰动类型,这种失败无法通过针对性修复解决,因为其触发条件从未出现。

英文摘要

CASCADE is an agentic framework that predicts downstream transcriptional effects of gene perturbation from precomputed ARACNe regulatory networks, exposed via MCP. Prior work validates such tools by checking whether predicted genes are known cancer genes (membership); we instead test whether the predicted direction of change matches reality, using focal-gene copy-number amplification as a dosage-based proxy for the inverse of knockdown against real TCGA patient tumor data. For MYC, CASCADE's predicted knockdown targets show strong concordance with real amplified-vs-non-amplified tumor expression across three cancer types (BRCA: 90.0%, COAD: 72.0%, STAD: 85.7%; all p<0.0013), well above permutation baselines, surviving a PAM50 subtype control and replicating in an independent cohort (METABRIC, 87.2%). Compared against curated MSigDB gene-set baselines via Fisher's exact test, CASCADE's accuracy is not shown to exceed existing public knowledge of MYC- or E2F-driven biology, though its gene-specific direction-calling clearly outperforms a naive uniform guess. Extending to fifteen additional genes, validation proves gene-specific rather than universal: proliferation-machinery regulators mostly replicate, while lineage-identity transcription factors and one cyclin-D paralog (CCND2) consistently fail, a pattern we discuss as a hedged, post-hoc hypothesis. We separately benchmark whether an LLM-based agent correctly grounds natural-language requests into CASCADE's real MCP tool calls. Across 35 queries, a documented local model reaches 71.4% exact match (85.7% for a larger model); schema and gene-alias failures are resolved by scale or server-side correction, but both models confidently default to the wrong perturbation type on ambiguous queries, a failure a targeted fix could not resolve because its trigger condition never occurs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑