arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SparseDesign:精确编码序列设计的规模化扩展

SparseDesign: Scaling Exact Coding-Sequence Design

Hao Lin, Jingjin Yu

arXiv 2609.32308首次发表:更新:

发表机构

Rutgers University(罗格斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SparseDesign 通过候选稀疏化加速精确编码序列设计,在保持等价性的同时大幅降低计算与内存开销,实现显著加速。

AI 中文摘要

在同义编码序列的精确优化中,联合折叠能量与密码子使用目标受限于昂贵的动态规划分割和庞大的工作集。SparseDesign 将候选稀疏化应用于 Turner 2004 dangle-0 求解器在加权密码子自动机上的多环递归。仅当直接分支在相同端点状态的每个可分割或端点未配对实现上严格改进时,才保留该分支。我们在显式标量分支接口假设下,证明了在实数运算中与稠密递归的等价性。对于具有 N 个自动机状态、边集 E 和 Z 个保留候选的多环工作量为 O(N^2+N|E|+NZ);对于有界宽度自动机,最坏情况时间仍为立方级,总内存仍为二次方。端点所有权允许无锁并行候选构建。虽然合成压力家族几乎不能从稀疏化中受益,并表现出接近二次方的候选增长,但天然蛋白质显示出显著的候选数量减少。在我们的 7,600 任务活动中,2,000 蛋白质的人类表组在 λ=0 时中位保留率仅为 3.53%,在 λ=4 时为 2.15%,相对于所有可行直接区间,分别对应约 28.3 倍和 46.4 倍的减少。主要性能实验使用 AMD EPYC 7313 服务器。对于人类 Dp427c(11,031 nt,λ=0),16 线程打包的 SparseDesign 实现了五次运行中位数 236.54 秒墙钟时间和 14.43 GiB 峰值 RSS。与同一服务器上的单线程本地稠密 LinearDesign 分支(4,912 秒,402.10 GiB RSS)相比,这实现了 20.8 倍墙钟加速和 27.9 倍峰值内存减少。在配备 64 GiB RAM 的 Core i9-14900KF 商用 PC 上,相同输入、布局和线程数实现了 126.42 秒和 14.43 GiB RSS。

英文摘要

Exact optimization of synonymous coding sequences under a joint folding-energy and codon-usage objective is limited by expensive dynamic-programming splits and large working sets. \textsc{SparseDesign} applies candidate sparsification to the multiloop recurrence of a Turner~2004 dangle-0 solver over a weighted codon automaton. A direct branch is retained only when it strictly improves on every partitionable or endpoint-unpaired realization of the same endpoint states. We prove equivalence to the dense recurrence in real arithmetic, under an explicit scalar branch-interface assumption. With $N$ automaton states, edge set $E$ and $Z$ retained candidates, multiloop work is $O(N^2+N|E|+NZ)$; worst-case time remains cubic for bounded-width automata and total memory remains quadratic. Endpoint ownership permits parallel candidate construction without locks. While synthetic stress families can benefit little from sparsification and exhibit near-quadratic candidate growth, natural proteins show substantial candidate-count reductions. In our 7,600-task campaign, the 2,000-protein human-table panel has median retention of only 3.53\% at $λ=0$ and 2.15\% at $λ=4$, corresponding to approximately 28.3-fold and 46.4-fold reductions relative to all feasible direct intervals. The primary performance experiments use an AMD EPYC 7313 server. For human Dp427c (11,031 nt, $λ=0$), 16-thread packed \textsc{SparseDesign} achieves five-run medians of 236.54 seconds wall-clock time and 14.43 GiB peak RSS. Compared with the single-thread local dense LinearDesign fork on the same server (4,912 seconds, 402.10 GiB RSS), this gives a 20.8-fold wall-clock speedup and a 27.9-fold peak-memory reduction. On a Core i9-14900KF commodity PC with 64 GiB RAM, the same input, layout and thread count achieve 126.42 seconds and 14.43 GiB RSS.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑