发表机构
Johannes Kepler University Linz(约翰开普勒林茨大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现推理后端是影响LLM基准测试分数的不可忽视因素,约39%的分数变异性源于后端,建议披露后端信息及完整生成配置以提升结果可靠性。
AI 中文摘要
基准测试分数被视为模型的属性,然而用于生成这些分数的推理框架,如HuggingFace、vLLM或Ollama,却被认为不具影响力,其名称和版本几乎从未被披露。本研究调查了这种选择对模型输出的影响程度。在一项完全交叉的研究中(3个指令微调模型×5个推理框架×6个基准测试×4种生成模式),我们探究了不同工具(包装器/后端)如何影响基准测试分数,以及生成超参数如何影响分数变化。我们发现后端是不可忽视的因素,即使在无采样噪声的贪心解码下,更换后端也会显著改变模型性能,且这种影响是结构性的、强烈依赖于模型的。按生成模式分解方差后发现,从业者开箱即用看到的相当一部分变异性(约39%)可能源于后端,剩余部分源于采样噪声和每个框架的默认生成参数,而通过披露和匹配生成配置可避免这两种情况。这种差异在事实类基准测试上比在社会偏见类基准测试上更明显。总体而言,基准测试数字并非与后端无关,因此我们建议披露后端、其版本以及完整的生成配置,同时在跨后端比较时使用确定性解码。
英文摘要
Benchmark scores are reported as properties of a model, yet the inference framework used to produce them, such as HuggingFace, vLLM, or Ollama, are considered non-influential and their names and versions are almost never disclosed. In this work we investigate how much this choice can influence the model output. In a fully-crossed study (three instruction-tuned models x five inference frameworks x six benchmarks x four generation modes) we investigate how different tools (wrappers/backend) influence benchmark scores and how their score changes is influenced by generation hyper-parameters. We find backend to be a non-negligible factor where even under greedy, sampling-noise-free decoding, changing the backend can significantly alter models performance and this effect is structural and strongly model-dependent. Decomposing the variance according to generation mode reveal that considerable portion of the variability (roughly 39\%) a practitioner sees out-of-the-box can stem from the backend, while the remaining stems from sampling noise and each framework's default generation parameters, both of which are avoidable by disclosing and matching the generation configuration. These divergences are more pronounced on factual than on social-bias benchmarks. Overall, benchmark numbers are not backend-agnostic therefore, we recommend disclosing the backend, its version, and the full generation configuration, also using deterministic decoding for cross-backend comparison.