arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11889cs.SEcs.AI

评估真实程序中的精确输出与检查点状态预测

Evaluating Exact Output and Checkpoint-State Prediction in Real Programs

Xiaohong Chen, David Bucur, Chenglong Ma, Yi Zhang, Lingming Zhang, Sriram Vishwanath, Grigore Rosu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出了用于预测真实程序最终输出与检查点状态的基准,评估了四类模型的七种设置,发现启用推理的设置表现更优,且基准可揭示短输出分数隐藏的错误。

中文摘要 AI 辅助

我们提出了一个仅从源代码和输入预测最终输出及检查点状态的基准。该基准扩展了CRUXEval风格的输出预测,包含配对的较短和较长轨迹输入,以及循环内部和之后的检查点。该基准包含来自371个Python和C++程序的400个案例,在四个模型家族的七种设置下进行评估,不使用工具或代码执行。在计划的11200次预测中,11151次产生了可分级的响应。启用推理的设置在已完成的响应上比其非推理对应设置高出33.1至55.2个百分点。最强设置在较短轨迹最终输出上得分为93.0%,较长轨迹最终输出上为77.0%,两项状态任务上分别为65.5%和63.5%;当缺失响应计为错误时,这些分数仍然成立。在2397个具有相同源代码的匹配Python比较中,将输入改为较长轨迹会产生528次正确到错误的变化和147次反转。按源代码聚类的分析保留了这一准确率差距,而调整后的Python模型未提供累积状态负载与错误之间存在正增量关联的证据。更改输入和检查点任务会共同改变多个因素,因此差距无法单独归因于轨迹长度或内部状态跟踪机制。该基准揭示了仅靠短输出分数所隐藏的错误。

英文摘要

We present a benchmark for predicting final output and checkpoint state from source and input alone. It extends CRUXEval-style output prediction with paired shorter- and longer-trace inputs and checkpoints inside and after a loop. The benchmark contains 400 cases from 371 Python and C++ programs, evaluated under seven settings from four model families without tools or code execution. Of 11,200 planned predictions, 11,151 produced gradable responses. Reasoning-enabled settings outperform their off counterparts by 33.1 to 55.2 percentage points on completed responses. The strongest setting scores 93.0% on shorter-trace final output, 77.0% on longer-trace final output, and 65.5% and 63.5% on the two state tasks; these scores also hold when missing responses count as wrong. Across 2,397 matched Python comparisons with identical source, changing to the longer-trace input yields 528 correct-to-wrong changes and 147 reversals. Source-clustered analyses preserve this accuracy gap, while adjusted Python models give no evidence of a positive incremental association between cumulative state load and error. Changed inputs and checkpoint tasks alter several factors together, so the gaps do not isolate trace length or an internal state-tracking mechanism. The benchmark exposes errors hidden by short-output scores alone.

发表机构

  • Intent Computing
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑