arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39644cs.LG

RiboUnmix:从有偏且含噪的Ribo-seq测量中学习共享翻译动态

RiboUnmix: Learning Shared Translational Dynamics from Biased and Noisy Ribo-seq Measurements

Gabriele Martino, Denis Skibinski, Ivo L. Hofacker, Sebastian Tschiatschek

首次发表
浏览论文内容

中文总结 AI 辅助

RiboUnmix提出一种概率多数据集框架,通过分离共享序列依赖信号与实验效应,从有偏含噪的Ribo-seq数据中恢复真实翻译动态,并在合成与真实基准上优于基线。

中文摘要 AI 辅助

核糖体图谱分析(Ribo-seq)测量沿mRNA的核糖体分布,但观察到的占据率图谱也包含实验特异性失真和随机变异性。因此,能准确预测测量图谱的模型可能复现技术效应,而非恢复潜在的生物学信息。我们探究是否联合建模在不同实验条件下收集的数据集,能够揭示共享的、序列依赖的核糖体占据模式。我们提出RiboUnmix,一个概率性多数据集框架,其中每个期望的测量图谱被表示为共享的序列依赖信号,该信号由数据集特异性的乘法因子调制。一个负二项观测模型捕捉了重复之间的变异性。我们在一个受控的合成基准上评估RiboUnmix,该基准结合了编程翻译动力学、核糖体交通、随机计数采样和序列依赖的实验失真。由于潜在的动力学和失真是已知的,共享图谱和数据集特异性效应的恢复可以分别评估。两个推断的成分与其目标高度相关,表明RiboUnmix能够将共享的动力学模式与实验效应区分开来。在四个物种特异性的真实数据基准上,RiboUnmix在预测测量图谱方面优于序列到图谱基线。在114个HEK来源数据集的子集上独立训练的模型,对于保留的转录本恢复了一致的共享图谱,并且改变训练数据集的数量和组成的实验表明,学习到的表示保持稳定。因此,RiboUnmix将实验间的变异转化为可重复的序列依赖核糖体占据模式的证据,支持从多样化的Ribo-seq数据集中生成生物学假设。

英文摘要

Ribosome profiling (Ribo-seq) measures ribosome distributions along mRNAs, but observed occupancy profiles also contain experiment-specific distortions and stochastic variability. Consequently, models that accurately predict measured profiles may reproduce technical effects rather than recover the underlying biology. We ask whether jointly modeling datasets collected under different experimental conditions can reveal shared, sequence-dependent patterns of ribosome occupancy. We introduce RiboUnmix, a probabilistic multi-dataset framework in which each expected measured profile is represented as a shared sequence-dependent signal modulated by a dataset-specific multiplicative factor. A negative-binomial observation model captures variability across replicates. We evaluate RiboUnmix on a controlled synthetic benchmark combining programmed translation kinetics, ribosome traffic, stochastic count sampling, and sequence-dependent experimental distortions. Because the underlying kinetics and distortions are known, recovery of the shared profile and dataset-specific effects can be assessed separately. Both inferred components correlate strongly with their targets, demonstrating that RiboUnmix can disentangle shared kinetic patterns from experimental effects. Across four organism-specific real-data benchmarks, RiboUnmix outperforms sequence-to-profile baselines in predicting measured profiles. Models trained independently on subsets of 114 HEK-derived datasets recover concordant shared profiles for held-out transcripts, and experiments varying the number and composition of training datasets show that the learned representation remains stable. RiboUnmix thus converts variation across experiments into evidence for reproducible sequence-dependent patterns of ribosome occupancy, supporting biological hypothesis generation from diverse Ribo-seq datasets.

发表机构

  • University of Vienna(维也纳大学)

机构由 AI 辅助整理,请以论文原文为准。

↑