arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05610cs.ETcs.LG

LC-Implicit-QAOA:用于有界QUBO光锥上训练的激活工作空间上限精确目标与梯度评估

LC-Implicit-QAOA: Active-Workspace-Capped Exact Objective-and-Gradient Evaluation for Training over Bounded QUBO Light Cones

  • Institute of Smart Industry and Green Energy, National Yang Ming Chiao Tung University(国立阳明交通大学智慧产业与绿能研究院)

机构由 AI 辅助整理,请以论文原文为准。

Chih-Chung Hsu

AI总结:

该研究提出LC-Implicit-QAOA算法,通过限制因果锥与伴随微分,在工作空间预算下高效完成有界QUBO光锥上的QAOA训练,大幅降低梯度评估的时间与内存开销。

AI中文摘要:

QAOA训练会反复查询目标函数及所有共享梯度,即便QUBO项具有有界因果锥,精确评估仍是可行性瓶颈。基于已确立的因果锥限制与伴随微分,LC-Implicit-QAOA在局部振幅与命名工作空间分配前,先分析锥结构与诱导边数量,随后在命名激活评估器工作空间预算下,联合选择等大小微批次与检查点调度。“Implicit”指省略全局状态与全局代价表,而非隐式微分;不可行请求会在分配前被拒绝。独立实现的complex128/float64稠密伴随在1800余次图角对比中与LC结果一致,最坏相对梯度误差为1.56×10^-13。LC在p=2的有界锥网格中完成全部104个目标请求;在预设n≤24的验证上限下,28个请求执行匹配的状态加代价参考,76个请求故意不运行。在80个预算内请求中,测得的已分配评估器内存始终在预算内,最高为预算的0.797倍。在3-正则n=512、p=2的设置下,伴随以101次目标等效调用、189秒达到相同有限预算终点,而中心差分法需909次调用、1565秒。LC针对固定深度的单局域与双局域对角QUBO代价,采用横向场混频器;它不提供全局状态、采样,也不提供与硬件无关的最快后端规则。

英文摘要:

QAOA training repeatedly queries an objective and all shared gradients, making exact evaluation a feasibility bottleneck even when QUBO terms have bounded causal cones. Building on established causal-cone restriction and adjoint differentiation, LC-Implicit-QAOA profiles cone structure and induced-edge counts before local-amplitude and named-workspace allocation, then jointly selects equal-size microbatches and checkpoint schedules under a named active-evaluator workspace budget. "Implicit" means omitting both global state and global cost table, not implicit differentiation; infeasible requests are rejected before those allocations. An independently implemented complex128/float64 dense adjoint agrees with LC over 1,800 graph-angle comparisons, with a worst relative gradient error of 1.56 x 10^-13. LC completes all 104 target requests in a p=2 bounded-cone grid; under a prespecified n <= 24 validation cap, the matched state-plus-cost reference is executed for 28 requests and deliberately not run on 76. Across 80 budgeted requests, measured allocated evaluator memory stays within budget, reaching at most 0.797 of it. On 3-regular n=512, p=2, the adjoint reaches the same finite-budget endpoint in 101 objective-equivalent calls and 189 s, versus 909 calls and 1,565 s for central differences. LC targets fixed-depth one- and two-local diagonal QUBO costs with a transverse-field mixer; it provides neither global states, sampling, nor a hardware-independent fastest-backend rule.

补充信息

↑