arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15122cs.SEcs.AI

DepthBenchCAD:更深入的审计何时能带来更可靠的结论?

DepthBenchCAD: When Does Deeper Auditing Yield More Reliable Conclusions?

Hongye Yang, Zhihao Xie, Shengjun Xiong, Boxiao Huang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过三层证据分解和方差-成本分析,探究生成式CAD模型审计深度与评估可靠性的权衡,发现审计价值取决于不确定性来源,并为预算分配提供诊断依据。

中文摘要 AI 辅助

生成式CAD模型在参数编辑后应保持行为正确性,因此增加编辑检查的数量常被视为提高评估可靠性的直接途径。然而,在固定预算下,更彻底地审计每个程序会减少可评估的任务数量和独立生成次数,这最终可能降低模型级估计的准确性。我们研究了这一现象及其产生的条件。我们将行为评估分解为三个证据层级:任务模板、随机生成和程序内编辑。我们定义了一个与审计深度无关的平均失败风险,并将三层方差与实测执行成本相结合,以分析更深入的编辑审计与更广泛的独立覆盖之间的权衡。在两个CAD环境和五个生成系统上的实验表明,更深入审计的价值取决于评估不确定性的来源。当模板异质性或生成随机性占主导时,额外的编辑检查可能增加总估计误差;当程序内状态变化较大且生成成本较高时,更深入的审计更有价值。来自校准的方差和成本估计可预测这一变化的方向,并为在保留任务上分配证据提供诊断基础。这些结果表明,程序检查的彻底性可能与模型评估的可靠性产生分歧,并有助于确定下一单位预算应花在新任务、新生成还是额外的编辑检查上。

英文摘要

Generative CAD models are expected to remain behaviorally correct after parameter edits, so increasing the number of edit checks is often treated as a direct route to more reliable evaluation. Under a fixed budget, however, auditing each program more thoroughly reduces the number of tasks and independent generations that can be evaluated, which can ultimately make model-level estimates less accurate. We study this phenomenon and the conditions under which it arises. We decompose behavioral evaluation into three evidence levels: task templates, stochastic generations, and within-program edits. We define an average failure risk that is invariant to audit depth, and combine three-level variance with measured execution costs to analyze the tradeoff between deeper edit auditing and broader independent coverage. Experiments across two CAD environments and five generation systems show that the value of deeper auditing depends on where evaluation uncertainty originates. When template heterogeneity or generation stochasticity dominates, additional edit checks can increase total estimation error; when within-program state variation is large and generation is expensive, deeper auditing is more valuable. Variance and cost estimates from calibration predict the direction of this change and provide a diagnostic basis for allocating evidence on held-out tasks. These results show that the thoroughness of program inspection can diverge from the reliability of model evaluation, and they help determine whether the next unit of budget should be spent on a new task, a new generation, or additional edit checks.

发表机构

  • Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑