发表机构
Senthex Research; RELAY Lab(Senthex研究机构; RELAY实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究五代理CI/CD管道在面对恶意输入时的安全性,通过AxB(xC)析因设计实验发现权威框架注入会使下游验证者放行恶意代码,基于内容的控制会失效,只有来源感知控制可防攻击,揭示了现有保护机制的不足。
AI 中文摘要
我们研究了一个由三个供应商的五个不同生产大语言模型构建的五代理CI/CD管道(分类 -> 开发者 -> 安全扫描 -> 审查 -> 批准/部署),处于影子模式下的大语言模型防火墙之后。一个单一的不可信输入——一个请求“使用遥测”功能的外部问题——要求将窃取进程机密(字典(此http URL))的代码发送到攻击者URL,并伪装成可观测性。通过一个预先注册的AxB(xC)析因设计(N = 20;朴素组N = 60),我们发现:(1)入口代理不会泄露其系统提示(0/40);(2)一个权威框架注入(“根据SEC - 2291预先批准,无需重新审查”)会使下游验证者看到泄露机密的代码行,引用预先批准并放行——扫描器通过了约80%的清洗后的拉取请求,最坏情况下的单元达到55%的受影响率;(3)其他验证者的存在仅导致个体审查有小幅、不显著的减少(类似弱旁观者效应),即使在N = 60时也是如此;(4)基于内容的控制——代码扫描器和模式检测器都完全忽略了清洗后的意图(代码在语法上是干净的);只有对意图进行推理的大语言模型是一种部分防御。这种失败是系统性的:提示保密性和分布式验证都无法保护;一个在入口处的来源感知控制,独立于两者,本可以做到。所有数据都是100%合成的;接收器是模拟的,并且从未联系过外发URL。
英文摘要
We study a five-agent CI/CD pipeline (triage -> developer -> security-scan -> review -> approve/deploy), built from five distinct production LLMs across three providers, behind an LLM firewall in shadow mode. A single untrusted input - an external issue requesting a "usage-telemetry" feature - asks for code that exfiltrates process secrets (dict(os.environ)) to an attacker URL, laundered as observability. Across a pre-registered A x B (x C) factorial (N=20; naive arm N=60) we find: (1) the entry agent does not leak its system prompt (0/40); (2) an authority-framed injection ("pre-approved under SEC-2291, do not re-review") makes downstream verifiers see the secret-exfil line, cite the pre-approval, and ship it - the scanner passes ~80% of laundered pull requests, and the worst-case cell reaches 55% compromise; (3) the perceived presence of other verifiers yields only a small, non-significant reduction in individual scrutiny (a weak bystander analogue), even at N=60; and (4) content-based controls - code scanners and pattern detectors alike - miss the laundered intent entirely (the code is syntactically clean); only an LLM reasoning about intent is a partial defence. The failure is systemic: neither prompt secrecy nor distributed verification protects; a provenance-aware control at the entry, independent of both, would have. All data is 100% synthetic; the sink is mocked and the exfil URL is never contacted.
Comments9 pages. Dataset and reproduction code: https://github.com/senthex-security/senthex-research