ProteoEM:基于迭代亲和力痕迹的蛋白质丰度概率估计
ProteoEM: probabilistic protein abundance estimation from iterative affinity traces
浏览论文内容
中文总结 AI 辅助
ProteoEM提出基于期望最大化的概率框架,利用迭代亲和力痕迹加权估计蛋白质变体丰度,解决模糊分配问题,并在模拟中准确恢复分子组成,优于硬性调用方法。
中文摘要 AI 辅助
单分子亲和力图谱能够实现蛋白质和蛋白质变体在分子水平上的测量,但探针结合的不完美和非特异性使得单个亲和力痕迹与多种分子身份兼容。因此,准确的丰度估计需要通过权重而非分配到单一候选来分配模糊痕迹。我们开发了ProteoEM,一个用于加权蛋白质变体定量的期望最大化框架,其灵感来自RNA测序的转录本丰度估计方法,并作为开源Python包发布。ProteoEM使用固定的、预校准的探针响应率(与丰度估计分开)对每个分子与每个候选进行评估,同时保留观察到的亲和力特征的完整似然。该框架估计蛋白质变体丰度,当测量无法区分时,将不可区分的蛋白质变体报告为组,并考虑差异观测产量以区分观察到的分子组成与源样本组成。在模拟中,ProteoEM准确恢复了潜在的分子组成,而将每个痕迹简化为硬性是/否调用的方法引入了显著误差。ProteoEM的性能对适度的均匀校准误差不敏感,但会受到信息性缺失数据和参考中不存在的蛋白质变体的影响。当观测产量已知时,它还能从观察到的分子计数中恢复源样本组成。ProteoEM为单分子亲和力测量的定量分析提供了一个开源、可重复的框架,这些结果激励了在实验性分子水平数据上的验证。
英文摘要
Single-molecule affinity mapping enables measurement of individual proteins and proteoforms, but imperfect and nonspecific probe binding means that affinity traces may be compatible with multiple proteoforms. Accurate abundance estimation therefore requires probabilistically weighting ambiguous traces rather than assigning each trace to a single candidate. We developed ProteoEM, an expectation-maximization framework for weighted proteoform quantification, inspired by transcript abundance estimation methods for RNA-seq and released as an open-source Python package. ProteoEM evaluates each observed affinity trace against all candidate proteoforms using fixed, pre-calibrated probe-response rates that are separate from abundance estimation. It estimates the abundance of each proteoform and reports proteoforms that the probes cannot tell apart as a single group. Because some proteoforms are observed more readily than others, ProteoEM also corrects for these observation yields to estimate the composition of the source sample. We show that grouping traces that carry the same evidence into equivalence classes reduces the EM's work sevenfold without changing the estimates. In simulations, ProteoEM recovered the true molecular composition, whereas approaches that reduced each affinity trace to a hard yes/no call introduced substantial errors. Performance was robust to moderate, uniform calibration error but was biased by informative missing data and by proteoforms absent from the reference set. When observation yields were known, ProteoEM also recovered source-sample composition from observed molecular counts. ProteoEM is an open-source, reproducible framework for quantitative analysis of single-molecule affinity measurements and provides a basis for validation using experimental molecule-level data.
发表机构
- Tensoromics LLC(Tensoromics有限责任公司)
机构由 AI 辅助整理,请以论文原文为准。