HarnessSecurity-Bench:安全机制真的能保护编码智能体框架吗?
HarnessSecurity-Bench: Do Security Mechanisms Really Protect Coding Agent Harnesses?
浏览论文内容
中文总结 AI 辅助
针对编码智能体框架安全机制实证不足的问题,提出HarnessSecurity-Bench基准,系统评估六框架九机制,发现自动批准显著提升攻击成功率,并给出安全设置可验证等改进建议。
中文摘要 AI 辅助
编码智能体框架(coding agent harnesses)负责中介工具使用并授权操作,但其安全机制和运行时影响尚未得到充分刻画。我们提出了 HarnessSecurity,这是首个针对开源和闭源编码智能体框架的系统性实证研究与基准测试。首先,我们推导出一个包含十种机制的分类体系,然后使用研究人员和大语言模型(LLM)评审员的独立评分,评估了 400 个“框架-机制”组合单元。我们发现,已确认的机制实现中约有一半是可选加入(opt-in)的,而闭源框架则存在明显的证据空白。其次,我们引入了 HarnessSecurity-Bench,这是一个包含 23 个任务、覆盖五个攻击面且不牺牲合法任务需求的基准测试。通过使用独立的确定性预言机(deterministic oracles)来测量任务效用和攻击效果,并结合安全设置比较,我们评估了六个领先框架中的九种机制:Claude Code、Codex CLI、Gemini CLI、gptme、Qwen Code 和 GitHub Copilot。在受控的 LLM 基线 GLM-5.2 下,我们进行了 2,500 次试验,记录了 81,155 次工具调用和超过 22 亿个令牌。启用自动批准(auto-approve)会提高效用,并将攻击成功率从 29.2% 提升至 95.6%。网络隔离和只读模式能降低攻击效果,但会带来显著的效用损失;而命令允许列表和命令拒绝列表在分别带来较小效用损失和效用提升的同时,降低了攻击效果。任务级案例表明,对共享能力的限制可能同时阻碍合法和恶意操作,并且允许的工具或命令可能通过替代执行路径使未经授权的操作仍可触达。框架提供商应使安全设置可验证,测试通往受保护操作的替代执行路径,并在评估任务效用和执行成本的同时评估攻击效果。
英文摘要
Coding agent harnesses mediate tool use and authorize actions, yet their security mechanisms and runtime effects remain incompletely characterized. We present HarnessSecurity, the first systematic empirical study and benchmark of open- and closed-source coding agent harnesses. First, we derive a ten-mechanism taxonomy and then assess 400 harness-mechanism cells using independent ratings by researchers and large language model (LLM) judges. We find that about half of confirmed mechanism implementations are opt-in, while closed-source harnesses exhibit substantial evidence gaps. Second, we introduce HarnessSecurity-Bench, a benchmark of 23 tasks across five attack surfaces without sacrificing legitimate task requirements. Using separate deterministic oracles to measure task utility and attack effects with security setting comparisons, we evaluate nine mechanisms across six leading harnesses: Claude Code, Codex CLI, Gemini CLI, gptme, Qwen Code, and GitHub Copilot. Under a controlled LLM baseline GLM-5.2, we conduct 2,500 trials, recording 81,155 tool calls and over 2.2 billion tokens. Enabling auto-approve increases utility and raises attack success from 29.2% to 95.6%. Network isolation and read-only mode reduce attack effects with substantial utility losses, while command allowlisting and command denylisting reduce attack effects with a small utility loss and a utility gain, respectively. Task-level cases show that restrictions on a shared capability can obstruct both legitimate and malicious operations, and that allowed tools or commands can leave unauthorized operations reachable through alternative execution paths. Harness providers should make security settings verifiable, test alternative execution paths to protected operations, and assess attack effects alongside task utility and execution costs.
发表机构
- School of Software Engineering, Sun Yat-sen University(中山大学软件学院)
- Peng Cheng Laboratory(鹏城实验室)
- East China Normal University(华东师范大学)
- Department of Computer Science, Hong Kong Baptist University(香港浸会大学计算机系)
- Tsinghua University(清华大学)
- Department of New Networks, Peng Cheng Laboratory(鹏城实验室新网络部)
机构由 AI 辅助整理,请以论文原文为准。