评估结论的可复现性如何?对LLM推断提示结构的一次自我审计
How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure
浏览论文内容
中文总结 AI 辅助
本研究通过自我审计LLM提示结构推断评估,发现小样本评估结论高度不稳定,排名仅底部可靠,建议报告排名稳定性与原始数据以增强可信度。
中文摘要 AI 辅助
对LLM系统的评估通常会在少量提示集上取平均值,并将模型呈现为排名表。我们以基于LLM的提示结构推断为案例研究,探讨这样的排名表究竟值得多大程度的信任:涉及五个家族的八个开源模型变体,参数规模从8B到675B,禁用缓存,并持久化了293个原始中间表示。所测量的现象本身就不稳定。相同的调用并不能可靠地恢复相同的结构,平均节点集Jaccard指数从0.39到0.96不等,且72%的提示-模型组合从未达到节点集完美匹配。对评估进行审计会进一步削弱其结论,而这是我们的主要贡献。在提示上的联合聚类自助法下,只有排名底部是稳固的:两个最不可复现的模型在99%和86%的重复样本中保持排名,中间四个模型在27%到48%之间,而前两个模型各为68%,因此该表能可靠地识别最差模型,但不能可靠地识别最佳模型。两种同样合理的合并重复实验的规则改变了八行中的四行,并将研究层面的头条指标移动了7个百分点。将推断结构与真实标注进行对比显示,可复现性不能等同于准确性。此外,八个端点中有四个在测量后十周内被撤回,因此所规定的研究已无法再运行。因此,小样本LLM评估可能看起来比其证据所支持的更具决定性。我们建议在报告任何排名时,同时报告排名稳定性、每个单元格的来源、已执行的敏感性比较、每次运行的原始输出以及测量日期。
英文摘要
Evaluations of LLM systems routinely average over small prompt sets and report models as a ranked table. We ask how much confidence such a table deserves, using LLM-based prompt-structure inference as the case study: eight open model variants across five families and 8B to 675B parameters, caching disabled, 293 raw intermediate representations persisted. The measured phenomenon is unstable to begin with. Identical calls do not reliably recover identical structure, with mean node-set Jaccard from 0.39 to 0.96 and 72% of prompt-model cells never node-set-perfect. Auditing the evaluation weakens its conclusions further, and this is our main contribution. Under a joint cluster bootstrap over prompts, only the bottom of the ranking is firm: the two least reproducible models hold rank in 99% and 86% of replicates, the middle four in 27% to 48%, and the top two in 68% each, so the table identifies the worst model reliably but does not reliably identify the best. Two equally defensible rules for merging repeated campaigns change four of eight rows and move the study-wide headline by 7 percentage points. Checking the inferred structure against ground-truth annotations shows reproducibility cannot be read as accuracy. And four of the eight endpoints were withdrawn within ten weeks of measurement, so the study as specified can no longer be run. Small-sample LLM evaluations can therefore look far more definitive than their evidence supports. We recommend reporting rank stability, per-cell provenance, executed sensitivity comparisons, raw per-run outputs, and a measurement date alongside any ranking.
发表机构
- Skelf Research
机构由 AI 辅助整理,请以论文原文为准。