arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PromptResponse:优化大语言模型编码任务的提示词

PromptResponse: Optimizing Prompts for LLM Coding Tasks

Erik Thureck, Robert Kühnen, Tim Jacobowitz

arXiv 2608.21074首次发表:更新:

发表机构

Humboldt-Universität zu Berlin; HU Berlin(柏林洪堡大学; 柏林洪堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出PromptResponse,通过对HumanEval数据集的5种变体让GPT-4o执行8200次编码任务,发现统一格式(尤其JSON)可提升代码效率与稳定性,LLM调优提示词会降低性能,还发布了相关数据集与评估流程。

AI 中文摘要

大型语言模型(LLM)越来越多地被用于研究工作流和软件开发流程中,但其输出仍对输入提示词的变化十分敏感。本文提出了《PromptResponse》,一项受控研究,探究编码任务提示词的格式及基于LLM的调优如何影响生成代码的性能、效率与稳定性。研究采用HumanEval数据集的5个语义相同但句法不同的变体——基线版、JSON版、Markdown版、YAML版及LLM调优版——让GPT-4o完成编码任务,共执行8200次。结果显示,统一格式(尤其是JSON格式)可提升生成效率与句法稳定性,任务性能也有小幅提升;而LLM调优后的提示词会导致任务性能显著下降,且在其他维度无明显改善。这些发现表明,仅低工作量的重新格式化即可产生可衡量的提升,而调优需考虑模型对齐问题。本文最后基于研究结果提供了一套实用建议,并发布了数据集变体与评估流程以供未来研究使用。

英文摘要

Large language models (LLMs) are increasingly used in research workflows and software development pipelines, yet their output remains sensitive to input prompt variations. This paper presents $\unicode{x00AB}$PromptResponse$\unicode{x00BB}$, a controlled study examining how formatting and LLM-based tuning of coding task prompts affect the resulting code's performance, efficiency, and stability. Using five semantically identical yet syntactically distinct variants of the HumanEval dataset$\unicode{x2014}$baseline, JSON, Markdown, YAML, and an LLM-tuned version$\unicode{x2014}$we had GPT-4o solve its coding problems over 8200$\unicode{x00A0}$executions. Our results show that consistent formatting$\unicode{x2014}$especially JSON$\unicode{x2014}$improves generation efficiency and syntactic stability, with minor gains in task performance. Conversely, the LLM-tuned prompts resulted in significantly degraded task performance without significant improvements in any other dimension. These findings suggest that low-effort reformatting alone can yield measurable improvements, while tuning must account for model alignment. We conclude our work with providing a set of practical recommendations informed by our results as well as releasing our dataset variants and evaluation pipeline for future work.

Comments22 pages, 7 figures, 10 listings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑