arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于不等式约束的无松弛深度展开组合优化求解器

Slack-Free Deep-Unfolded Combinatorial Optimization Solver for Inequality Constraints

Ryo Hagiwara, Shunta Arai, Satoshi Takabe

arXiv 2607.20042首次发表:更新:

AI 中文总结

研究有不等式约束的组合优化问题求解,提出结合不平衡惩罚与大胁方法的UPOM及深度展开的DU-UPOM,通过辅助变量更新等改进,减轻调整和嵌入负担,在随机背包问题实验中,DU-UPOM能在更少迭代达最优解。

AI 中文摘要

量子退火用于解决组合优化问题,在量子退火器上实现时通常编码为二次无约束二元优化问题,但约束编码会增加量子比特数和嵌入开销。对于有不等式约束的组合优化问题,标准松弛变量公式会引入额外二元变量。不平衡惩罚(UP)避免了松弛变量,但原始UP公式需调整两个惩罚系数且含平方残差项。本文提出不平衡惩罚大胁方法(UPOM),将UP与大胁方法结合用于不等式约束组合优化问题。UPOM用辅助变量更新取代原始UP的两个静态惩罚系数,并从用于采样的哈密顿量中去除平方残差项。还提出深度展开不平衡惩罚大胁方法(DU-UPOM),从训练实例中学习UPOM更新的步长调度。随机背包问题的数值实验表明,UPOM优于原始UP,DU-UPOM比固定步长UPOM和其他基线在更少迭代中达到最优解。这些结果表明,所提框架减轻了UP的调整和嵌入负担,同时使大胁方法可用于不等式约束训练。

英文摘要

Quantum annealing (QA) is used to solve combinatorial optimization problems (COPs). When COPs are implemented on quantum annealers, they are typically encoded as quadratic unconstrained binary optimization (QUBO) problems, but constraint encodings often increase the number of qubits and the embedding overhead. This issue is particularly important for COPs with inequality constraints, where standard slack-variable formulations introduce additional binary variables. Unbalanced penalization (UP) avoids slack variables, but the original UP formulation requires tuning two penalty coefficients and contains a squared residual term that can increase the number of quadratic couplings. In this paper, we propose the unbalanced penalization Ohzeki method (UPOM), which combines UP with the Ohzeki method for inequality-constrained COPs. UPOM replaces the two static penalty coefficients of original UP with an auxiliary-variable update and removes the squared residual term from the Hamiltonian used for sampling. We further propose the deep-unfolded unbalanced penalization Ohzeki method (DU-UPOM), which learns the step-size schedule of the UPOM update from training instances. Numerical experiments on random knapsack problems show that UPOM improves over original UP and that DU-UPOM reaches optimal solutions in fewer iterations than fixed-step UPOM and other baseline. These results demonstrate that the proposed framework reduces the tuning and embedding burdens of UP while making the Ohzeki method trainable for inequality constraints.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑