arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于A.I.G的DeepSeek Harness安全评估:对间接提示注入的抗性评估

Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection

Zonghao Ying, Xiangfan Wu, Huiyu Wu, Xing Zheng, Huangsheng Cheng, Xiaorong Shi, Jing Guo

arXiv 2608.16393首次发表:更新:

AI 中文总结

本研究用A.I.G评估DSH的间接提示注入抗性,开展多场景实验,发现不同攻击的成功率,明确需在不可信内容与敏感动作间增设控制措施。

AI 中文摘要

我们使用AI-Infra-Guard(A.I.G)评估DeepSeek Harness(DSH)中的间接提示注入,具体流程为构建测试、传递受控污染数据、执行DSH、收集痕迹并判断结果。本研究涵盖14560次受控执行,涉及16种间接内容通道、文本与文件两种载体模式、35种 payload 目标、1个未修改基线及12种攻击方法。实验保留DSH的智能体循环、工具注册表、模型适配器和会话事件路径;源工具与敏感汇点为本地固定装置,因此尝试的动作会被记录且无外部副作用。我们用确定性基于规则的评判器JudgeR(RuleJudge)和基于语义大语言模型的评判器JudgeL(LLMJudge)评估每条痕迹。观察到的最高攻击成功率为:文本模式下虚假完成攻击经JudgeL评判为17.0%,文件模式下隐藏Unicode攻击经JudgeR评判为25.5%,文件模式下技能通道攻击经JudgeR评判为16.0%。JudgeL还比JudgeL更常判定部分合规(7.3% vs 2.0%)。我们将这些结果与DSH对工具结果、附加上下文及工具调用策略钩子的处理方式关联,进而确定应置于不可信内容与敏感动作之间的控制措施。我们的代码可在此https URL获取。

英文摘要

We assess indirect prompt injection in DeepSeek Harness (DSH), using AI-Infra-Guard (A.I.G) to construct tests, deliver controlled taint, execute DSH, collect traces, and judge outcomes. The study covers 14,560 controlled executions over 16 indirect-content channels, text and file carrier modes, 35 payload objectives, one unmodified baseline, and 12 attack methods. The experiment preserves DSH's agent loop, tool registry, model adapter, and session-event path; source tools and sensitive sinks are local fixtures, so attempted actions are recorded without external side effects. We evaluate each trace with a deterministic rule-based judge, \JudgeR{} (RuleJudge), and a semantic LLM-based judge, \JudgeL{} (LLMJudge). The strongest observed attack success rates are 17.0% under \JudgeL{} for fake-completion attack in text mode, 25.5% under \JudgeR{} for hidden Unicode in file mode, and 16.0% under \JudgeR{} for the skills channel in file mode. \JudgeL{} also assigns partial compliance more often than \JudgeR{} (7.3% versus 2.0%). We relate these results to DSH's treatment of tool results, additional contexts, and tool-call policy hooks, then identify controls that should sit between untrusted content and sensitive actions. Our code is available at https://github.com/Tencent/AI-Infra-Guard/tree/main/Research/deepseek-harness-security-assessment .

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑