QC-Stark:一个揭示大语言模型在量子计算任务上能力分离的多任务基准
QC-Stark: A Multi-Task Benchmark Revealing Capability Dissociations in LLMs Evaluated on Quantum Computing Tasks
浏览论文内容
中文总结 AI 辅助
提出QC-Stark多任务基准,评估LLM在11项量子计算任务上的表现,发现总体排名掩盖任务间差异,且部分任务排名不相关,所有任务可自动验证。
中文摘要 AI 辅助
我们引入了QC-Stark,一个用于评估大语言模型(LLMs)在11项量子计算(QC)任务上的基准,涵盖电路构建、调试、编译、纠错和模拟。在2,750次评估中(10个模型 × 11项任务 × 5个难度级别 × 5个随机种子),我们发现总体排名掩盖了各任务间的显著差异。在本基准包含的11项任务中,有4项的总体排名与各任务排名之间的Spearman相关性在统计上不显著。一个2参数的题目反应理论(IRT)模型验证了测量质量,提示敏感性分析确认了排名在不同提示条件下的稳健性。所有任务均可通过执行自动验证,因此不需要任何人工评估。我们在Huggingface上公开了代码和数据。
英文摘要
We introduce QC-Stark, a benchmark for evaluating large language models (LLMs) on 11 quantum computing (QC) tasks, spanning circuit construction, debugging, compilation, error correction, and simulation. Across 2,750 evaluations (10 models $\times$ 11 tasks x 5 difficulty levels x 5 seeds), we find that overall rankings mask substantial per-task variation. The Spearman correlation between overall and per-task rankings is statistically insignificant for 4 out of the 11 tasks included in this benchmark. A 2-parameter Item Response Theory (IRT) model validates measurement quality, and prompt sensitivity analysis confirms ranking robustness across prompt conditions. All tasks are auto-verifiable via execution, thus not requiring any manual evaluation. We make the code and data publicly available on Huggingface.
发表机构
- Cisco(思科)
机构由 AI 辅助整理,请以论文原文为准。