发表机构
Adobe Inc(Adobe公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提示包装器格式差异对模型分数的影响,引入格式敏感性指数(FSI)和可解析性敏感性指数(PSI),通过大量实验发现平均FSI在不同模型间变化超30倍,且可解析性是准确性有力预测指标,指出不考虑包装器差异和合规性报告准确性不可靠并给出建议。
AI 中文摘要
提示包装器通常仅在格式上有所不同,但它们可以使模型分数变化到足以改变排行榜结论。我们在令牌控制协议下研究这种差异,并引入两个互补指标:格式敏感性指数(FSI),即由包装器选择引起的准确性范围;以及可解析性敏感性指数(PSI),即答案可解析性的相应范围。在跨越7个问答任务、5个包装器家族以及4个参数从7B到72B的指令模型的140,000次OpenRouter生成中,我们发现平均FSI在不同模型间变化超过30倍,且很大程度上由合规失败导致。固定效应回归表明,即使在控制任务、模型和包装器后,可解析性仍是准确性的有力预测指标。我们认为,在不报告包装器差异和合规性的情况下报告准确性在统计上是不可靠的,并为基准测试和结构化输出部署给出了实际建议。
英文摘要
Prompt wrappers often differ only in formatting, yet they can change model scores enough to flip leaderboard conclusions. We study this variance under a token-controlled protocol and introduce two complementary metrics: the Format Sensitivity Index (FSI), the accuracy range induced by wrapper choice, and the Parseability Sensitivity Index (PSI), the corresponding range in answer parseability. Across 140,000 OpenRouter generations spanning 7 QA tasks, 5 wrapper families, and 4 instruct models from 7B to 72B parameters, we find that mean FSI varies by over 30x across models and is largely explained by compliance failures. A fixed-effects regression shows that parseability remains a strong predictor of accuracy even after controlling for task, model, and wrapper. We argue that reporting accuracy without wrapper variance and compliance is statistically fragile, and we give practical recommendations for both benchmarking and structured-output deployments.
Comments10 pages, 6 figures