arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27234cs.LGq-bio.QM

发现、证伪、修订:从源代码到预测贡献审计智能体发现细胞模型中的输入使用声明

Discover, Falsify, Revise: Auditing Input-Use Claims from Source Code to Predictive Contribution in Agent-Discovered Cell Models

Mengran Li, Bo Li, Chengyang Zhang, Yang Yan, Jinfeng Xu, Zhenchao Tang

中文总结 AI 辅助

针对AI虚拟细胞中智能体模型对输入使用声明的不可靠问题,提出CELLAUDIT审计方法,通过源代码检查和预测贡献验证,结合证伪修订,提升模型性能与输入使用可靠性。

中文摘要 AI 辅助

AI虚拟细胞旨在预测细胞对特定干预的响应,然而仅凭留出预测性能并不能确立对所提供的扰动信息的使用。这一预测-声明差距在智能体模型发现中尤为重要,其中语言模型智能体使用基于评分的反馈生成并修订预测器。我们引入CELLAUDIT,通过询问输入是否能进入引用的计算、拟合的预测是否依赖于它、以及这种依赖性是否改善了对观测响应的预测,来审计输入使用声明。在配对的形态-转录组扰动基准(BBBC047)上,智能体选择的预测器达到平均留出全局皮尔逊相关系数(PCC)0.3153,但对化合物替换保持不变;仅对照谱的预测器达到0.3142。源代码检查识别出被单例键值注意力阻断的化合物查询路径,且在使用不相交对照孔重新拟合后不变性仍然存在。在对两个关联任务的48个候选者的分层审计中,47个在化合物替换下在两个留出折上改变预测,但只有20个在两个折上显示目标损失增益且区间高于零。在BBBC047上,证伪引导的修订恢复了正的化合物平均贡献,同时保持相对于仅对照谱基线的增益。在匹配的sci-Plex搜索中,审计富集的反馈产生更高的留出性能和跨五个轨迹的更大的平均化合物和剂量贡献,尽管配对区间跨越零。在独立获取的队列上重新拟合固定设计表明,预测泛化不一定意味着输入使用声明的泛化:剂量贡献持续存在,而化合物身份的支持则不然。CELLAUDIT为智能体模型发现增加了一个证伪层,从生成-评分-修订转向发现-证伪-修订。

英文摘要

AI virtual cells aim to predict cellular responses to specified interventions, yet held-out predictive performance alone does not establish use of the supplied perturbation information. This prediction-claim gap matters in agentic model discovery, where language-model agents generate and revise predictors using score-based feedback. We introduce CELLAUDIT, which audits input-use claims by asking whether an input can enter the cited computation, whether fitted predictions depend on it, and whether that dependence improves prediction of observed response. On a paired morphology-transcriptomics perturbation benchmark (BBBC047), an agent-selected predictor attains a mean held-out Global Pearson correlation coefficient (PCC) of 0.3153 but remains invariant to compound replacement; a control-profile-only predictor reaches 0.3142. Source inspection identifies a compound-query pathway blocked by singleton key-value attention, and the invariance persists after refitting with disjoint control wells. In a stratified audit of 48 candidates across two linked tasks, 47 change predictions under compound replacement on both held-out folds, but only 20 show target-loss gains with intervals above zero on both folds. On BBBC047, falsification-guided revisions recover positive mean compound contributions while retaining gains over the control-profile-only baseline. In matched sci-Plex searches, audit-enriched feedback yields higher held-out performance and larger mean compound and dose contributions across five trajectories, although paired intervals span zero. Refitting fixed designs on an independently acquired cohort shows predictive generalization need not imply generalization of input-use claims: dose contribution persists, whereas support for compound identity does not. CELLAUDIT adds a falsification layer to agentic model discovery, moving from generate-score-revise toward discover-falsify-revise.

↑