arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20319quant-phcs.LG

QEncodeBench:大型语言模型能否将经典问题编码为经过验证的量子预言机?

QEncodeBench: Can Large Language Models Encode Classical Problems into Verified Quantum Oracles?

  • University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校)
  • George Mason University(乔治梅森大学)
  • University of North Texas(北德克萨斯大学)

机构由 AI 辅助整理,请以论文原文为准。

Xujun Che, Hanhan Wu, Yuchen Yuan, Chenyang Yu

AI总结:

QEncodeBench评估LLM将经典约束问题编码为量子相位预言机的能力,发现语义错误主导且推理模式显著提升准确率,神经符号流水线通过所有测试。

AI中文摘要:

Grover搜索、振幅放大和量子计数都依赖于同一个可复用的子程序,即相位预言机,算法文献假定其构造是现成的:经典谓词被假定为已经编码为正确的、资源有界的电路。我们将这一假定转化为一项可测量的能力。QEncodeBench要求大型语言模型(LLMs)将经典约束问题编码为相位预言机,并使用对抗性自验证的验证器对生成的电路进行评分,该验证器决定到全局相位为止的完整解集等价性,同时恢复辅助比特并强制执行资源预算。我们表明,采样的基态测试会系统性地高估这种能力。以此方式衡量,模型之间出现明显分化:没有推理模式的代码模型几乎无法解决任何问题,而在相同权重上启用原生推理可将准确率提高一个数量级。失败主要在于语义层面而非语法层面。两种架构,即单元验证的约束智能体和神经符号编译流水线,通过将正确性关键的组合委托给确定性程序,弥补了剩余差距的大部分。消融实验量化了每个组件的贡献,资源门控揭示了依赖于架构的电路宽度与深度之间的权衡。最后,受控的难度升级揭示了架构对难度结构的特定响应:不同的难度轴会降低不同方法的性能,而神经符号流水线通过了所有评估实例。代码和数据可在以下网址获取:https://this https URL。

英文摘要:

Grover search, amplitude amplification, and quantum counting all rely on the same reusable subroutine, a phase oracle, whose construction the algorithms literature takes as given: the classical predicate is assumed to be already encoded as a correct, resource-bounded circuit. We turn this assumption into a measured capability. QEncodeBench tasks large language models (LLMs) with encoding classical constraint problems as phase oracles and scores the generated circuits with an adversarially self-validated verifier that decides full solution-set equivalence up to a global phase, with ancillas restored and resource budgets enforced. Sampled basis-state tests, we show, systematically overestimate this ability. Measured this way, models separate sharply: code models without a reasoning mode solve essentially nothing, and enabling native reasoning on identical weights improves accuracy by an order of magnitude. The failures are overwhelmingly semantic rather than syntactic. Two architectures, a unit-verified constraint agent and a neuro-symbolic compilation pipeline, close most of the remaining gap by delegating correctness-critical composition to deterministic procedures. Ablations quantify the contribution of each component, and resource gating exposes an architecture-dependent trade-off between circuit width and depth. Finally, controlled difficulty escalation reveals architecture-specific responses to difficulty structure: different difficulty axes degrade different methods, while the neuro-symbolic pipeline passes every evaluated instance. Code and data are available at https://github.com/chexujun/QEncodeBench.

↑