发表机构
Shanghai Artificial Intelligence Laboratory; Shanghai Jiao Tong University(上海人工智能实验室; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大语言模型推理出错时现有交互方法的问题,提出深度交互机制,可直接编辑原始响应并提炼精炼提示引导模型,在STEM任务推理中纠正成功率大幅提升,令牌使用量显著减少。
AI 中文摘要
思维链(CoT)推理的出现显著增强了大语言模型(LLMs)处理复杂多步任务的能力。然而,当出现错误时,当前交互方法通常涉及重新生成可能再次出错的另一个响应,或者用户费力地在后续轮次中标记错误步骤,这可能会得到类似“你是对的,我在这里犯了错误”的回复,随后类似错误会再次出现。为解决此问题,我们提出一种用于精确纠正LLMs推理错误的高效人工干预机制,称为深度交互。我们的方法能够直接编辑原始响应,在保留准确推理步骤的同时纠正错误部分。我们将编辑后的CoT提炼为一个精炼提示,然后引导LLM沿着纠正后的推理路径进行。实验结果表明,与基线方法相比,我们的方法在STEM任务推理中的纠正成功率提高了25%以上,令牌使用量减少了约40%。
英文摘要
The emergence of Chain-of-Thought (CoT) reasoning has significantly enhanced the ability of large language models (LLMs) to tackle complex, multi-step tasks. However, when errors occur, current interaction approaches typically involve re-generating another response that may make mistakes again, or users laboriously flag the faulty step in follow-up turns that may get responses <You are right, I made a mistake here> followed by similar errors recurring. To address this issue, we propose an efficient human intervention mechanism for precisely correcting reasoning errors in LLMs, termed Deep Interaction. Our approach enables direct editing of the original response, allowing erroneous parts to be corrected while preserving accurate reasoning steps. We refine the edited CoT into a distilled prompt, which then steers the LLM along the corrected reasoning path. Experimental results show that our method achieves over a 25% improvement in correction success rate and reduces token usage by approximately 40% on STEM tasks reasoning compared to baseline approaches.