arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于非靶向 LC-MS 代谢组学中可重复特征选择的多宇宙共识管道

A multiverse-consensus pipeline for reproducible feature selection in untargeted LC-MS metabolomics

Mohammed Saeed Al-Huraibi, Ihsan Yozgat, Ahmet Kaplan

arXiv 2607.17345首次发表:更新:

发表机构

Istanbul Medipol University; Istinye University(伊斯坦布尔梅迪波尔大学; 伊斯坦布尔伊斯蒂尼耶大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对非靶向 LC-MS 代谢组学预处理决策多且结果依赖未知选择的问题,提出多宇宙共识管道,通过十阶段质量控制及多方法组合分析,在示范数据集上得出更稳健的特征选择结果,还讨论了方法的范围与局限。

AI 中文摘要

背景:非靶向 LC-MS 代谢组学需要一系列预处理决策,每个决策都有多种合理选项。分析人员通常采用一种管道并报告最终的特征候选列表。但该列表对未改变的选择的依赖程度却不为人知。结果:我们将多宇宙分析应用于非靶向代谢组学特征选择。提出了一个可审计、由配置驱动的管道,它应用十阶段质量控制过滤器级联并记录每个特征的命运,还针对四种对比预处理理念运行下游分析,每种理念与四种特征排序方法结合,在自举稳定性选择和标签置换测试下进行。只有在各路径中重复出现的特征才进入分层共识。在五个乳腺癌细胞系的示范数据集上,四个单管道分别返回 4 - 20 个特征的候选列表,其成对一致性低至 Jaccard = 0.05。多宇宙共识保留了 15 个特征,其中一个在所有四条路径中都出现,尽管两条路径(共享归一化和漂移校正方法)主导了共识。全管道标签置换测试在 50 次空置换中未发现错误发现。结论:仅报告预处理稳健的特征并带有完整的保留/丢弃审计跟踪,将隐藏的分析自由度转化为明确、可检查的输出。我们讨论了范围和局限性,包括单批次设计和独立验证的必要性。

英文摘要

Background: Untargeted LC-MS metabolomics requires a long chain of preprocessing decisions, each with several equally defensible options. Analysts typically commit to one pipeline and report the resulting feature shortlist. How strongly that shortlist depends on choices that were never varied stays invisible. Results: We adapt multiverse analysis to untargeted metabolomics feature selection. We present an auditable, configuration-driven pipeline that (i) applies a ten-stage quality-control filter cascade in which every feature's fate is logged, and (ii) runs the downstream analysis as a multiverse over four contrasting preprocessing philosophies, each combined with four feature-ranking methods under bootstrap stability selection and label-permutation testing. Only features recurring across paths enter a tiered consensus. On a demonstration dataset of five breast-cancer cell lines (30,370 detected features), the four single pipelines individually returned shortlists of 4-20 features whose pairwise agreement was as low as Jaccard = 0.05. The multiverse consensus retained 15 features (>=2/4 paths), of which one recurred across all four, although two paths (sharing normalization and drift-correction methods) dominate the consensus. A pipeline-wide label-permutation test found no false discoveries in 50 null permutations. Conclusions: Reporting only preprocessing-robust features, with a complete kept/dropped audit trail, converts hidden analytical degrees of freedom into an explicit, inspectable output. We discuss scope and limitations, including single-batch design and the need for independent validation.

Comments14 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑