arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10057quant-phcs.AIcs.CVcs.LG

量子电路视觉:用于量子代码生成的视觉人工智能代理的成本感知评估

Quantum Circuit Vision: Cost-Aware Evaluation of Visual AI Agents for Quantum Code Generation

Dongping Liu, Aoyu Zhang, Luyao Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

研究人工智能代理对量子电路图的理解及代码生成成本,提出量子电路视觉评估框架,构建基准并评估模型,发现中级模型在成本-准确性上平衡最佳,电路深度是失败主因,提出级联路由策略并开源数据集及代码。

中文摘要 AI 辅助

人工智能代理能否直观理解量子电路图并生成经过验证的可执行代码,且成本如何?我们提出了量子电路视觉,这是一个用于多模态人工智能代理量子电路视觉理解的成本感知评估框架。我们构建了一个包含13个类别的132个电路的基准(1至10个量子比特),带有可执行的亚马逊Braket代码和酉保真度验证。通过对三个不同能力成本层级的前沿Claude系列模型进行n = 5次重复试验评估,我们发现中级模型(Sonnet 4.6,成本为最强模型Opus 4.6每次调用成本的18%)在成本-准确性前沿提供了最有利的平衡:核心子集的通过率为91%,最强模型的准确性优势在统计上不显著(配对t检验:p = 0.083)。逻辑回归证实电路深度而非量子比特数是失败的主要预测因素(p < 0.001)。思维链提示没有统计学上的显著效果(所有p > 0.18,n = 5),这表明对于结构耦合图,视觉模式识别比明确的推理策略更重要。我们提出了一种级联路由策略(从便宜到昂贵的模型),在单模型成本的38%时实现了84%的准确率,表明模型路由作为一种成本杠杆比提示工程更重要。我们在Hugging Face Hub上发布了QCV - 数据集(132个电路,5种模态,1931个文件)作为开放评估基础设施,带有结构化元数据以实现可发现性、互操作性和负责任的人工智能文档,并在GitHub上提供所有评估代码、成本日志和验证脚本以实现完全可重复性。

英文摘要

Can AI agents visually comprehend quantum circuit diagrams and generate verified executable code--and at what cost? We present Quantum Circuit Vision, a cost-aware evaluation framework for multimodal AI agents on quantum circuit visual understanding. We construct a 132-circuit benchmark spanning 13 categories ($1$--$10$ qubits) with executable Amazon Braket code and unitary-fidelity verification. Evaluating three frontier Claude-family models at different capability-cost tiers with $n=5$ repeated trials, we find that the mid-tier model (Sonnet 4.6, $1.30\times$ credits) offers the most favorable balance on the cost-accuracy frontier: 91% pass rate on the core subset at 18% of the per-call cost of the strongest model (Opus 4.6), whose accuracy advantage is not statistically significant (paired $t$: $p=0.083$). Logistic regression confirms that circuit depth--not qubit count--is the primary predictor of failure ($p<0.001$). Chain-of-thought prompting shows no statistically significant effect (all $p>0.18$, $n=5$), suggesting that visual pattern recognition outweighs explicit reasoning strategy for structurally coupled diagrams. We propose a cascade routing strategy (cheap $\rightarrow$ expensive models) that achieves 84% accuracy at 38% of single-model cost, demonstrating that model routing dominates prompt engineering as a cost lever. We release QCV-Dataset (132 circuits, 5 modalities, 1,931 files) on Hugging Face Hub as an open evaluation infrastructure with structured metadata for discoverability, interoperability, and responsible AI documentation, and all evaluation code, cost logs, and verification scripts on GitHub for full reproducibility.

发表机构

  • Amazon Web Services(亚马逊网络服务)
  • Duke Kunshan University(杜克昆山大学)

机构由 AI 辅助整理,请以论文原文为准。

↑