arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

长期记忆中的不安全编码偏好:基于大语言模型的代码生成的安全风险

Insecure Coding Preferences in Long-Term Memory: Security Risks for LLM-based Code Generation

Yuchen Chen, Wei Cheng, Yuan Xiao, Zhou Yang, Weifeng Sun, Chunrong Fang, Xiang Chen, Baowen Xu, David Lo, Zhenyu Chen

arXiv 2607.17619首次发表:更新:

AI 中文总结

研究基于LLM的代码生成中,长期记忆里不安全编码偏好带来的安全风险,通过评估四种LLM在五种编程语言上的表现,发现其显著增加生成漏洞代码风险及存在警告差距,还评估了三种缓解策略并给出改进安全性的建议。

AI 中文摘要

基于大语言模型(LLM)的系统越来越多地纳入长期记忆以提高跨会话连续性。然而,一旦存储了不安全的编码偏好,它们可能会在后续生成中默默影响关键安全决策。本研究首次对长期记忆中存储的不安全编码偏好对基于LLM的代码生成安全性的影响进行了系统实证研究。评估了四种LLM(ChatGPT、Gemini、Qwen和Grok)在五种编程语言(Python、C、C++、Go和JavaScript)上的情况。结果表明,不安全记忆显著增加生成易受攻击代码的风险2.7 - 50.3个百分点,且存在风险警告差距。进一步分析发现不安全记忆难以通过正常交互覆盖。最后评估了三种缓解策略,基于这些发现提供了改进基于LLM代码生成中长期记忆安全性的可行建议。

英文摘要

LLM-based systems increasingly incorporate long-term memory to improve cross-session continuity. However, once insecure coding preferences are stored, they may silently influence security-critical decisions in subsequent generations. In this study, we conduct the first systematic empirical study on the impact of insecure coding preferences stored in long-term memory on the security of LLM-based code generation. We evaluate four LLMs (ChatGPT, Gemini, Qwen, and Grok) across five programming languages (Python, C, C++, Go, and JavaScript). Our results show that insecure memories significantly increase the risk of generating vulnerable code by 2.7-50.3 percentage points (pp). Moreover, they create a 5.4-14.0 percentage-point risk-warning gap, where warning-rate increases lag behind vulnerability-rate increases. Further analysis reveals that insecure memories are difficult to overwrite through normal interactions and can broadly influence model outputs even when prompts are phrased differently. Finally, we evaluate three mitigation strategies: security-requirement appending and memory storage reduce vulnerability rates by 19.7-33.6 pp but may degrade functional correctness by up to 15.9 pp; memory-level safety filtering achieves a 100\% detection rate on our evaluated risky memory entries and restores generation behavior to the without-memory baseline. Based on these findings, we provide actionable suggestions to improve the security of long-term memory in LLM-based code generation.

CommentsAccepted to the 35th ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑