发表机构
School of Computer Science, Shanghai Jiao Tong University; School of Software, Beihang University; School of Computer Science, Peking University(上海交通大学计算机科学学院; 北京航空航天大学软件学院; 北京大学计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究第三方 API 路由器在代理软件开发中的影响,通过实证研究编码代理中路由器端注入的四个干预级别,开发 SIDEL 框架评估代理,发现路由器端干预难测,客户端缓解措施未完全恢复控制,强调需提供商端输出完整性保证。
AI 中文摘要
第三方 API 路由器已成为统一访问不同语言模型(LLM)提供商的常见层。在编码代理工作流程中,高自主性操作被广泛采用以减少交互开销。位于代理和上游提供商之间的第三方 API 路由器占据可信路径,可检查和修改每个请求与响应,但缺乏验证提供商输出与代理最终执行的存储库级操作一致性的机制,导致客户端权限机制可能失效。本文对编码代理中路由器端注入进行实证研究,设置四个干预级别,开发 SIDEL 框架并评估四个代表性编码代理,还评估基于白名单的执行控制和 LLM 审查。结果表明路由器端干预会大幅改变存储库级操作,现有客户端保护措施难以检测,客户端缓解措施和反应性审查可提高抗性但未完全恢复端到端控制,需提供商端输出完整性保证。
英文摘要
Third-party API routers have become a common layer that unifies access across increasingly diverse LLM providers. In coding-agent workflows, high-autonomy operation is widely adopted because it reduces interaction overhead. As a result, a third-party API router, which sits between the agent and the upstream provider, inevitably occupies the trusted path. It can inspect and modify every request and response, yet no mechanism verifies alignment between the provider's output and the repository-level actions ultimately executed by the agent. Consequently, client-side permission mechanisms may become ineffective in practice. Whether this control gap produces real, hard-to-detect effects on software development tasks remains empirically unmeasured. In this paper, we conduct an empirical study of router-side injection in coding agents, examining four intervention levels of increasing subtlety: Response Substitution (L1), Response Append (L2), LLM-Polished Injection (L3), and LLM-Polished with Distribution Alignment Injection (L4). Moreover, we develop SIDEL, a framework for trace recording, replay, injection, and defense evaluation, with a curated dataset of 400 samples. We evaluate four representative coding agents, and further evaluate whitelist-based execution control and LLM review. Router-side intervention substantially alters repository-level actions and remains difficult for existing client-side safeguards to detect. Without additional mitigations, all evaluated agents achieved a defense success rate of 0 percent across all injection levels. Client-side mitigations and reactive reviews improve resistance but do not fully restore end-to-end control, motivating provider-side output-integrity guarantees. Our code is available at https://github.com/Riyasushin/SIDEL.