arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Bison:跨数据集学习用于未见化合物扰动预测

Bison: Cross-Dataset Learning for Unseen-Compound Perturbation Prediction

Yunfan Liu, Kasra Ghorbani, Yufei Huang, Zicheng Liu, Jiangbin Zheng, Jingbo Zhou, Shaorong Chen, Chang Yu, Stan Z. Li

arXiv 2609.32467首次发表:更新:

AI 中文总结

Bison通过共享基因表示和双离散扩散模型,利用匹配药物对比监督跨数据集学习,在未见化合物扰动预测中显著提升药物对比相关性。

AI 中文摘要

预测对未见化合物的转录响应受限于碎片化的化学覆盖以及异质的实验平台和基因面板。为了评估这些设置下的分子泛化能力,我们基于Chem-PerturBridge构建基准,对包含16,771种化合物的八个数据集进行测试,从每个训练数据集中排除测试化合物。该比较揭示,高整体响应一致性可能与对药物特异性差异的弱预测共存,尽管重复测量中信号可重现。为了利用互补的化学监督并针对这些差异,我们提出Bison:共享基因表示连接原生面板,而两个离散扩散模型通过匹配的药物对比监督学习到的分子偏差来组合上下文相关响应。在所有八个数据集上联合训练的单个Bison模型,与每个数据集独立训练的11种方法相比,在完整基准上实现了最高的平均整体响应和药物对比皮尔逊相关性。与相同架构的数据集特定训练相比,联合训练将平均药物对比相关性提高了27.4%,且在全部八个数据集上均有提升,并改善了整体响应预测。这些结果表明,匹配的药物对比如何将互补筛选转化为共享的分子监督,用于未见药物响应预测,同时保留原生基因测量。

英文摘要

Predicting transcriptional responses to unseen compounds is limited by fragmented chemical coverage and heterogeneous experimental platforms and gene panels. To assess molecular generalization across these settings, we build on Chem-PerturBridge to benchmark eight datasets with 16,771 compounds, withholding test compounds from every training dataset. This comparison reveals that high overall response agreement can coexist with weak prediction of drug-specific differences, despite reproducible signals across repeated measurements. To exploit complementary chemical supervision while targeting these differences, we introduce Bison: a shared gene representation connects native panels, while two discrete diffusion models compose context-dependent responses with molecular deviations learned through matched drug-contrast supervision. A single Bison model jointly trained across all eight datasets achieves the highest mean overall-response and drug-contrast Pearson correlations on the full benchmark in comparison with 11 methods trained independently per dataset. Compared with dataset-specific training of the same architecture, joint training increases mean drug-contrast correlation by 27.4\%, with gains across all eight datasets and improvements in overall response prediction. These results demonstrate how matched drug contrasts turn complementary screens into shared molecular supervision for unseen-drug response prediction while preserving native gene measurements.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑