AI 中文总结
研究提示语气对大语言模型答案准确性及推理成本的影响,通过在MMLU数据集上对七种语气进行实验,发现输出令牌长度变化超准确性变化,不同模型在不同语气下有不同表现,揭示语气对答案质量和推理资源消耗的作用。
AI 中文摘要
我们研究了提示语气如何影响大语言模型答案的准确性以及以输出令牌消耗反映的推理成本。在一个包含570个问题的MMLU数据集上进行实验,以了解在七种不同语气(从谄媚到威胁)下大语言模型的准确性和推理成本之间的权衡。结果表明,所有模型的输出令牌长度变化大大超过准确性变化。输出令牌消耗在不同语气条件下最多相差44.3%。我们还分析了答案准确性与推理过程中平均输出令牌长度之间的权衡。对于ChatGPT模型4o和5 - 纳米,粗鲁语气占主导。对于Gemini模型2.5 Flash和2.5 Flash Lite,粗鲁和中性语气在帕累托最优前沿占主导。我们发现提示语气不仅影响答案质量,还影响现代大语言模型消耗的可计费推理资源量。
英文摘要
We examine how prompt tone affects both accuracy of the LLM answers and inference cost as reflected in output-token consumption. Experiments were performed to understand the trade-offs between accuracy and inference cost on a 570 Question MMLU dataset for LLM models prompted in seven different tones from sycophantic to threatening. Our results show that the output-token-length variation substantially exceeded accuracy variation across all models. Output-token consumption varied by up to 44.3% across tone conditions. We also analyzed the tradeoff between the accuracy of the answers and the average output token length in the reasoning process. For the ChatGPT models 4o and 5-nano, the rude tone is quite dominant. For the Gemini models 2.5 Flash and 2.5 Flash Lite, the rude and neutral tones are dominant on the Pareto-optimal frontier. We find that prompt tone influences not only answer quality but also the amount of billable inference resources consumed by modern LLMs.
Comments16 pages, 1 figure, 9-page Appendix