发表机构
University of Wisconsin–Milwaukee(威斯康星大学密尔沃基分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出六维提示侧结构复杂度指数,在代码生成前评分,发现通过率在综合评分13.75处存在非单调断点,任务类型和构造框架会移动断点,提供生成前可靠性测量框架。
AI 中文摘要
从生成代码中测得的复杂度依赖于失败情况:一个困难的提示可能产生一个短小的失败程序,并被赋予较低的输出复杂度。我们引入了一个六维度的提示侧结构复杂度指数,在生成之前进行评分,并与正确性分开考量。我们根据一个初步的单评分者评分标准,在六个区间中选取了5000个Python提示。四位外部大语言模型评分者对锁定的提示进行重新评分,得到19997行评分数据;在拥有全部四次评分的4998个提示上,综合评分者间信度为ICC = 0.872。我们对每个提示评估21个模型,共产生105000个生成结果。在未经调整的均值汇总分析中,通过率在综合评分13.75处存在非单调的断点,低于或等于该值时通过率为79.9%,高于该值时通过率为87.6%。这并非一个普遍的失败阈值。任务类型固定效应将断点移至10.75,并将区间差距从7.6个百分点缩小至2.1个百分点。一个构造框架控制将断点移至8.50,原始差距为-3.5个百分点,且单独任一框架都无法复现汇总的+7.6个百分点变化。模型特定拟合包括16个上升和5个下降的变化。一个包含365个提示的审计清理扩展在区间15和16上与原始五个模型的估计相匹配,但仅在区间16以上增加了14个提示。在具有可计算Lizard复杂度的零通过生成中,28.5%的生成将提示综合评分高于8与输出复杂度至多为10配对。在分歧增强的校准集上,人类一致性为中等且依赖于评分者;释义和跨语言重新评分保持了评分顺序。过度识别检验拒绝了对六个维度的联合限制,因此我们将该综合评分视为一个指数,不对2SLS估计做出因果解释。本研究的贡献在于一个生成前测量框架和对可靠性区间的有界观测分析。
英文摘要
Complexity measured from generated code is failure-dependent: a difficult prompt can yield a short failing program and be assigned low output complexity. We introduce a six-dimension prompt-side structural-complexity index scored before generation and kept separate from correctness. We select 5,000 Python prompts across six bands of a preliminary single-rater rubric. Four out-of-panel LLM raters rescore the locked prompts, giving 19,997 score rows; composite inter-rater reliability is ICC = 0.872 on the 4,998 prompts with all four ratings. We evaluate 21 models per prompt, yielding 105,000 generations. In the unadjusted mean-pooled analysis, pass rate has a nonmonotone breakpoint at composite 13.75, with 79.9% at or below and 87.6% above. This is not a universal failure cutoff. Task-type fixed effects shift the breakpoint to 10.75 and cut the regime gap from 7.6 to 2.1 points. A construction-frame control shifts it to 8.50 with a raw gap of -3.5 points, and neither frame alone reproduces the pooled +7.6-point change. Model-specific fits include 16 upward and five downward changes. A 365-prompt audit-clean extension matches the original five-model estimates at bins 15 and 16 but adds only 14 prompts above bin 16. Among zero-pass generations with computable Lizard complexity, 28.5% pair a prompt composite above 8 with output complexity at most 10. Human agreement is moderate and rater-dependent on a disagreement-enriched calibration set; paraphrase and cross-language rescoring preserve score ordering. Overidentification tests reject the joint restrictions on the six dimensions, so we treat the composite as an index and make no causal interpretation of the 2SLS estimates. The contribution is a pre-generation measurement framework and a bounded observational analysis of reliability regimes.
CommentsNeurIPS 2026 Evaluations & Datasets Track