arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DualityCert:量子场论中对偶性声明修复的验证器门控语言模型

DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory

Xingyang Yu

arXiv 2607.23614首次发表:更新:

发表机构

Virginia Tech(弗吉尼亚理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对量子场论中对偶性声明修复,提出DualityCert验证器。语言模型代理接收错误声明编辑至通过验证,在预注册基准测试中,验证器门控重试等策略在不同模型上有不同表现,各获胜策略都用低成本证书,还发布了相关内容。

AI 中文摘要

我们提出了DualityCert,这是一个用于四维N = 1箭袋规范理论中候选Seiberg对偶性声明的符号验证器。该验证器评估't Hooft反常匹配、超势R电荷一致性、中心电荷匹配和有界手征环代理。通过的声明会收到一致性证书,表明未发现测试的不一致性,但并非证明了对偶性。我们将验证器用作语言模型代理的修复环境,代理接收故意错误的声明并必须对其进行编辑直至通过验证。在145个错误声明的预注册基准测试中,验证器门控重试在deepseek-chat上比单次尝试提高最终修复成功率8.3个百分点,在qwen-plus上提高7.1个百分点(霍尔姆调整p < 0.002)。在十一次尝试的相同预算下,深度优先策略组合在deepseek-chat上比独立验证器过滤重采样表现差10.3个百分点,但在qwen-plus上比其好14.7个百分点,在两个验证模型中颠倒了两种测试的验证器利用策略的排序。在qwen-plus上,类别级验证器反馈比无内容重试价值高8.7个百分点,仅可解释的义务身份比结构相同的掩码反馈价值高6.4个百分点。在deepseek-chat上未检测到这些效果。另外,预注册的MiniMax-M2.5扩展再次发现迭代增益,独立验证器过滤重采样优于策略组合。因此,两种模型的更好策略不同,而每个获胜策略都使用相同的低成本证书。验证器、基准测试、协议和所有每次尝试的记录均已发布。

英文摘要

We present DualityCert, a symbolic verifier for candidate Seiberg-duality claims in four-dimensional N=1 quiver gauge theories. The verifier evaluates 't Hooft anomaly matching, superpotential R-charge consistency, central-charge matching, and a bounded chiral-ring proxy. A claim that passes receives a consistency certificate, which states that no tested inconsistency was found, not that the duality is proven. We use the verifier as a repair environment for language-model agents, which receive a deliberately broken claim and must edit it until it certifies. On a preregistered benchmark of 145 broken claims, with the analysis fixed before the first confirmatory model call, verifier-gated retry improves final repair success over a single attempt by +8.3 percentage points (pp) on deepseek-chat and +7.1 pp on qwen-plus (Holm-adjusted p<0.002). Under an equal budget of eleven attempts, the stop-first strategy portfolio underperforms independent verifier-filtered resampling by 10.3 percentage points on deepseek-chat but outperforms it by 14.7 points on qwen-plus, reversing the ordering of the two tested verifier-exploitation policies across the two confirmatory models. On qwen-plus, category-level verifier feedback is worth +8.7 pp over content-free retry, and interpretable obligation identities alone are worth +6.4 pp over structurally identical masked feedback. Neither effect is detected on deepseek-chat. Separately, a preregistered MiniMax-M2.5 extension again finds an iteration gain and independent verifier-filtered resampling outperforming the strategy portfolio. Which policy is better thus differs between the two models, while every winning policy uses the same cheap certificate. The verifier, benchmark, protocol, and all per-attempt records are released.

Comments17 pages, 2 figures, 9 tables. v2: added reference and note on concurrent related work. v3: added references. Code, benchmark, and all per-attempt records: https://github.com/xingyang-yu/QFTCert

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑