arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.09872q-bio.QM

生物序列分析中动态规划的微分框架

From Derivatives to Exact Sequence Substitution Effects in Dynamic Programming for Biological Sequence Analysis

Kiyoshi Asai

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对生物序列分析中动态规划问题,将多种模型表示为和积动态规划,定义相关量与序列变化,通过实验得出不同模型下的计算结果,此框架能确定导数适用情况及何时需重组,为相关分析提供统一基础。

中文摘要 AI 辅助

背景:生物序列分析中的动态规划通过对指数级数量的潜在路径、比对、推导树或RNA二级结构求和来计算概率或配分函数。其向后和外部量用于特定模型,但微分灵敏度与精确有限序列变化之间的关系很少在通用框架中阐述。方法:我们将隐马尔可夫模型、仿射间隙比对集合、随机上下文无关文法和RNA二级结构集合表示为和积动态规划,将向后和外部量定义为向前或内部变量的伴随,将序列变化定义为序列依赖局部因子的有限替换。结果:后验项边缘在内外积中归一化,局部事件后验还包括局部因子和子内部项,预期特征计数是配分函数的对数导数。对于隐马尔可夫模型、普通随机上下文无关文法和仿射间隙比对中的单位置替换,配分函数在位置特定因子组中是多仿射的,因此一位点变化可从一阶导数系数精确恢复,多位点变化可从混合导数恢复。在最近邻RNA模型中,替换会改变重叠环、堆积和多环因子以及边界上下文,因此精确的突变效应需要上下文依赖的内外重组,如Rchange算法。数值实验将暴力重新计算重现到机器精度。结论:该框架确定了何时导数给出精确的有限序列效应,何时需要更广泛的重组,为后验边缘、预期计数、参数灵敏度、突变分析和序列设计提供了统一基础。

英文摘要

Background: Dynamic programming in biological sequence analysis computes probabilities or partition functions by summing over exponentially many latent paths, alignments, derivation trees, or RNA secondary structures. Their backward and outside quantities are used model-specifically, but the relation between differential sensitivities and exact finite sequence changes is rarely stated in a common framework. Methods: We represent hidden Markov models, affine-gap alignment ensembles, stochastic context-free grammars, and RNA secondary-structure ensembles as sum--product dynamic programs, defining backward and outside quantities as adjoints of forward or inside variables and sequence changes as finite replacements of sequence-dependent local factors. Results: Posterior item marginals are normalized inside--outside products, local-event posteriors additionally include the local factor and child inside terms, and expected feature counts are logarithmic derivatives of the partition function. For HMMs, ordinary SCFGs, and single-position substitutions in affine-gap alignment, the partition function is multi-affine in position-specific factor groups, so a one-site change is recovered exactly from first-derivative coefficients and multisite changes from mixed derivatives. In nearest-neighbor RNA models a substitution alters overlapping loop, stacking, and multiloop factors and boundary contexts, so exact mutation effects instead require context-dependent inside--outside recombination, as in the Rchange algorithm. Numerical experiments reproduce brute-force recomputation to machine precision. Conclusions: The framework identifies when derivatives give exact finite sequence effects and when broader recombination is required, providing a unified basis for posterior marginals, expected counts, parameter sensitivity, mutation analysis, and sequence design.

补充信息

↑