使用证据引导推理提示减轻代码异味检测中的大语言模型谄媚现象
Mitigating LLM Sycophancy in Code Smell Detection Using Evidence-Guided Reasoning Prompts
浏览论文内容
中文总结 AI 辅助
研究基于LLM的代码异味检测中谄媚偏差问题,通过MLCQ数据集评估不同提示框架影响,发现模型对提示变化敏感。提出证据引导去偏提示策略,降低决策不稳定性,提高鲁棒性,为解决该问题提供有效方法。
中文摘要 AI 辅助
大语言模型(LLMs)因其解释程序语义的能力,越来越多地用于代码异味检测任务。然而,其在该场景下的可靠性仍未得到充分探索,尤其是在不同的提示条件下,模型预测可能受外部线索而非代码特征影响。其中一个限制是谄媚偏差,即模型倾向于使其输出与用户提供的假设一致,而非进行客观分析。本文首次对基于LLM的代码异味检测中的谄媚偏差进行了系统实证研究。使用MLCQ数据集,评估了不同提示框架(如确认偏差、矛盾提示和错误前提)如何影响模型预测。结果表明LLMs对提示变化高度敏感,决策翻转率高达72%,错误对齐率超过90%。为解决此问题,提出了证据引导去偏提示(EGDP),一种结构化提示策略,可降低决策不稳定性并提高鲁棒性,将决策翻转率降至低至12%,错误对齐率降至低至21%,同时增加对基于结构的证据的依赖。研究结果表明,谄媚偏差对基于LLM的代码异味检测的可靠性构成了关键威胁,证据引导推理提供了一种有效且可推广的缓解方法。
英文摘要
Large Language Models (LLMs) are increasingly used for code smell detection tasks due to their ability to interpret program semantics. However, their reliability in this context remains poorly explored, particularly under varying prompt conditions where model predictions may be influenced by external cues rather than code characteristics. One such limitation is sycophancy bias, where models tend to align their outputs with user-provided assumptions instead of performing objective analysis. In this paper, we present the first systematic empirical study of sycophancy bias in LLM-based code smell detection. Using the MLCQ dataset, we evaluate how different prompt framings like confirmation bias, contradictory hints, and false premises affect model predictions. Our results show that LLMs are highly sensitive to prompt variations, with Decision Flip Rates reaching up to 72% and False Alignment Rates exceeding 90%, indicating substantial instability and agreement with misleading prompts. To address this issue, we propose Evidence-Guided Debiasing Prompting (EGDP), a structured prompting strategy that enforces evidence-first reasoning. EGDP reduces decision instability and improves robustness, lowering Decision Flip Rates to as low as 12% and False Alignment Rates to as low as 21%, while increasing reliance on structurally grounded evidence. Our findings demonstrate that sycophancy bias poses a critical threat to the reliability of LLM-based code smell detection, and that evidence-guided reasoning provides an effective and generalizable mitigation approach.
发表机构
- Institute of Information Technology(信息科技研究所)
- University of Dhaka(达卡大学)
机构由 AI 辅助整理,请以论文原文为准。