arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

将推理痕迹提炼为软件工程任务的建议提示词

Distilling Reasoning Traces into Advisory Prompts for Software Engineering Tasks

Faizan Faisal, Prem Devanbu, Toufique Ahmed

arXiv 2608.00437首次发表:更新:

发表机构

University of California, Davis; IBM(加州大学戴维斯分校; 国际商业机器公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文受编程学习过程启发,提出将LLM推理痕迹提炼为建议提示词的方法,可减少LLM代码错误,部分提示词还能跨模型迁移,同时明确了方法适用场景。

AI 中文摘要

语言模型被广泛用于生成代码及相关处理(如识别代码幻觉、可能的输入或预测输出),但大型语言模型(LLM)可能出现严重错误。核心问题在于,模型训练所用的代码仍以人类编写为主,存在缺陷,且难以找到完全无错误的足够大规模代码语料库,因此需要无需额外训练的推理时方法来减少LLM错误。混合推理模型提供的可切换“推理”或“思考”模式确实能降低错误,但推理会消耗额外资源。本文探究能否在不总是承担推理成本的情况下实现更好性能。编程学习者通过四个步骤避免错误:识别错误、反思导致错误的认知偏差(即“仔细思考”错误)、从反思中推断一般规则或经验、将这些经验内化为规则,导师指导的教程中常见这种苏格拉底式互动,可内化的规则示例包括“编码前重述需求以明确内容”。受此过程启发,本文提出一种方法:首先识别(低资源)LLM的“思考模式”可避免错误的示例,再用更大的LLM检查这些错误及同一LLM通过“思考”避免错误的过程,生成总结性解释,最后由大型LLM将这些解释总结为简短的建议提示词。该方法适用于许多中等规模模型,部分情况下,学到的“建议提示词”还可有效迁移至其他模型。本文还研究了语言模型产生的代码错误的性质,并表征了该方法适用的场景。

英文摘要

Language models are widely used for generating and otherwise processing code (e.g., identifying code hallucinations, possible inputs, or predicting outputs); however, LLMs can make mistakes, which can be serious. One key issue is that models are trained on (still) largely human-written, and thus imperfect, code; it's not easy to find sufficiently large code corpora that are entirely free of bugs. Thus, other inference-time ways of reducing LLM errors, without additional training, are desirable. "Reasoning" or "thinking" modes, exposed as a togglable feature by hybrid reasoning models, do reduce errors; however, reasoning consumes additional resources. This paper asks if better performance can be achieved without always incurring the cost of reasoning. Human students of programming learn to avoid mistakes by (a) identifying them, (b) reflecting upon the cognitive lapses that led to them (essentially, "thinking through" the errors), (c) inferring general rules or lessons from these reflections, and (d) internalizing these lessons into rules. In tutorial sessions with an instructor, this is a common Socratic interaction. Examples of such internalizable rules might include the nugget "Before coding, restate the requirements to clarify them." Inspired by this process, this paper describes an approach where we first identify examples in which "thinking mode" in a (low-resource) LLM avoids errors. These errors, and their avoidance via "thinking" in the same LLM, are then examined by a bigger LLM to generate summary explanations; these are then summarized by a large LLM into brief advisory prompts. This approach works on many modest-sized models; in some cases, the "advisory prompts" thus learned can also be gainfully transferred to other models. We also present investigations into the nature of coding errors that language models make, and a characterization of when this approach can be helpful.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑