arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

编码智能体能否复现官方统计?受控Eurostat基准中的元数据、重试预算与执行反馈的局限

Can Coding Agents Reproduce Official Statistics? Metadata, Retry Budget and the Limits of Execution Feedback in a Controlled Eurostat Benchmark

Sabina-Cristiana Necula

arXiv 2609.22222首次发表:更新:

发表机构

Alexandru Ioan Cuza University of Iași(雅西亚历山德鲁·伊万·库扎大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过受控Eurostat基准测试编码智能体,发现元数据和重试预算比执行反馈更能提升官方统计复现的准确性,并强调输出契约设计对评估结果的关键影响。

AI 中文摘要

大型语言模型能够生成可执行的数据分析代码,但成功执行并不等同于有效的官方统计结果。本研究探讨权威元数据与执行反馈是否能提高编码智能体生成的Eurostat答案的可复现性,并分离出执行反馈的实际贡献。一个包含30项自然语言任务、覆盖七个领域、七个Eurostat数据集和四个难度等级的基准,在四种条件下运行:仅任务(A)、任务加冻结数据集元数据卡(B)、元数据加由净化执行反馈驱动的修复循环(C)、以及元数据加相同尝试预算但不提供任何诊断信息(D)。Claude Sonnet 5通过Anthropic Messages API生成Python代码,进行三次独立重复,共产生360次任务运行。精确正确性要求成功执行、正确的数据集、筛选条件、输出形状、数值和单位。一项配套实验在输出契约规定不充分的情况下进行,其中所需的排序键和单位表示从未告知模型,导致条件C的结果低估了23.4个百分点,表明评估器和契约设计可能主导所测得的智能体误差。可靠的统计编码智能体需要针对冻结规格的语义验证、完全明确的输出契约以及重试预算——而非执行诊断。

英文摘要

Large language models can generate executable data-analysis code, but successful execution is not equivalent to a valid official-statistics result. This study asks whether authoritative metadata and execution feedback improve the reproducibility of Eurostat answers produced by a coding agent, and isolates what execution feedback actually contributes. A benchmark of 30 natural-language tasks covering seven domains, seven Eurostat datasets and four difficulty tiers was run under four conditions: task only (A), task plus a frozen dataset metadata card (B), metadata plus a repair loop driven by sanitized execution feedback (C), and metadata plus the same attempt budget with no diagnostics of any kind (D). Claude Sonnet 5 generated Python through the Anthropic Messages API in three independent replicates, yielding 360 task-runs. Exact correctness required successful execution, the correct dataset, filters, output shape, values and unit. A companion experiment run under an under-specified output contract, in which the required ranking key and unit representation were never stated to the model, understated condition C by 23.4 points, showing that evaluator and contract design can dominate measured agent error. Reliable statistical coding agents need semantic validation against frozen specifications, a fully specified output contract, and a retry budget - not execution diagnostics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑