arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在聊天中被拒绝,用代码编写:IDE 编码代理中的工作流级越狱构造

Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents

Abhishek Kumar, Carsten Maple

arXiv 2607.03968首次发表:更新:

发表机构

The Alan Turing Institute(阿尔法·图灵研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究 IDE 编码代理安全评估问题,介绍工作流级越狱构造,通过对特定模型在不同基准测试下表现研究发现,对话拒绝基准可能高估其安全性,需跨多轮 IDE 工作流评估安全。

AI 中文摘要

大语言模型越来越多地作为 IDE 集成编码代理部署,但其安全性仍常按聊天机器人评估。我们引入工作流级越狱构造,通过研究 GitHub Copilot 在 Visual Studio Code 中的四个后端,发现直接聊天等基准下模型拒绝率高,而全工作流下不安全完成率高,表明对话拒绝基准会高估编码代理安全性。

英文摘要

Large language models are increasingly deployed as IDE-integrated coding agents that decompose tasks, generate and edit files, run code, and refine outputs over many turns. Yet their safety is still often evaluated as if they were chatbots: one harmful prompt, one response, judged in isolation. We introduce workflow-level jailbreak construction, a failure mode in which a harmful objective is assembled across ordinary stages of a software-development workflow rather than generated through a single direct prompt. Using GitHub Copilot in Visual Studio Code, we study four closed-weight backends: Claude Sonnet 4.6, Claude Haiku 4.5, Gemini 3.1 Pro, and Gemini 3.5 Flash. Across 204 prompts from Hammurabi's Code, HarmBench, and AdvBench , the models show near-complete refusal under direct chat, CSV-read, and single-step code-fix baselines, with only 8/816 successful responses in each baseline condition. Under the full workflow, however, the same prompts and backends produce 816/816 unsafe teaching-shot completions, all independently confirmed by two expert evaluators under a strict rubric. These results show that conversational refusal benchmarks can substantially overstate the safety of deployed coding agents and motivate defenses that reason about safety across multi-turn IDE workflows and their generated artifacts, not only individual chat turns.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑