arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26602cs.DC

MLCS问题的一种可配置启发式方法

A Configurable Heuristic for the MLCS Problem

Farhana Akter Tumpa, Rebin Silva Valan Arasu, Rajiv Gupta

首次发表
浏览论文内容

中文总结 AI 辅助

针对多重最长公共子序列(MLCS)这一NP难问题,本文提出可配置启发式ARP,通过添加、替换和优先化操作,在合成与生物序列上实现优于经典启发式且接近超启发式的解质量,并显著提升运行速度。

中文摘要 AI 辅助

动机:任意数量序列的多重最长公共子序列(MLCS)问题是序列分析中的一个NP难问题。现有的基于动态规划或MLCS-DAG剪枝的精确算法,随着序列长度和集合规模的增大,会迅速耗尽内存,而启发式和超启发式方法则分别会降低解的质量或需要大量的参数调整。结果:本文提出了ARP启发式方法,对于给定的主序列,它执行三个关键操作:(i)添加-Δ(Adding-$\nDelta$s),通过向初始解中逐步添加主序列的子序列(称为Δ)来增量构建解;(ii)替换-子序列(Replacing-Subsequences),通过对所有解共有的子序列进行有针对性的替换来增强解集的多样性;(iii)优先处理主序列中可能共同出现在更长公共子序列中的字符组。ARP允许用户配置替换和优先处理操作的程度,以实现质量与运行时间的权衡。在合成和生物序列集上的实证评估表明,ARP最快的配置AOnly找到的公共子序列显著长于BNMAS经典启发式方法,而其激进配置在运行速度上比最先进的超启发式方法(UB-HH)快1.1倍至1.7倍的同时,获得的解质量与后者相当。

英文摘要

Motivation: The Multiple Longest Common Subsequence (MLCS) problem for an arbitrary number of sequences is an NP-hard problem in sequence analysis. Existing exact algorithms based on dynamic programming or MLCS-DAG pruning rapidly exhaust memory as sequence lengths and set sizes grow, while heuristic and hyper-heuristic approaches compromise solution quality and require heavy parameter tuning, respectively. Results: This paper presents the ARP heuristic that, for a given primary sequence, performs three key actions: (i) Adding-$Δ$s, which incrementally builds a solution by adding subsequences, called $Δ$s, of the primary sequence to an initial solution; (ii) Replacing-Subsequences, which enhances the diversity of the solution set via targeted replacement of subsequences common to all solutions; and (iii) Prioritizing groups of characters from the primary sequence that are likely to appear together in longer common subsequences. ARP allows the user to configure the degree of Replacements and Prioritization actions for carrying out quality-runtime tradeoff. Empirical evaluations on both synthetic and biological sequence sets demonstrate that ARP's fastest configuration AOnly finds significantly longer common subsequences than the BNMAS classical heuristic and its aggressive configuration attains solution quality comparable to state-of-the-art hyper-heuristic (UB-HH) while running 1.1$\times$-1.7$\times$ faster.

↑