AI 中文总结
针对AI简历筛选中的间接提示注入,提出ResumeShield开源防御与基准,通过通道分离和过滤堆栈实现零误报、高检测率,并揭示攻击者的隐藏困境。
AI 中文摘要
AI简历筛选器会读取被评估者提供的文档,从而颠倒了评估者与其所评估材料之间通常的信任关系。求职者利用这一点,通过白色文本、零号字体、隐藏元素、标记注释、文档元数据或零宽字符在简历中隐藏指令。人类审查员看不到任何异常,而天真的提取流水线会将隐藏文本放入模型提示中,使其被当作指令并被执行。这就是间接提示注入,被OWASP列为LLM01:2025,最近的测量研究报告称,在生产筛选语料库中,约有百分之一的简历存在此类攻击。我们提出了ResumeShield,一个开源防御方案和基准。该防御结合了三个过滤阶段与第四个架构阶段,后者将候选内容置于一个明确隔离的数据通道中,操作者的可信指令声明该通道为惰性。该基准构建了一个种子合成语料库,涵盖九种隐藏技术和两种载荷家族,一种使用有记录的措辞,另一种模拟自适应攻击者,其围绕过滤器进行改写,并且仅当筛选结果发生变化时才将攻击计为成功。在包含104份文档的语料库上,天真的流水线在每个注入案例中都被操纵,而受防御的流水线从未被操纵。检测达到了1.000的精确率和0.944的召回率,在干净简历上没有误报。消融研究表明,仅通道分离就消除了所有测量的攻击成功,而完整过滤堆栈在没有分离的情况下仍使16.7%的攻击有效。我们还发现了一个隐藏困境:每个逃避检测的载荷都是攻击者留下的可见载荷,从而放弃了驱动攻击的隐形性。ResumeShield以Apache 2.0许可证发布,仅包含合成数据。
英文摘要
An AI resume screener reads a document supplied by the person it is evaluating, inverting the usual trust relationship between an assessor and the material it assesses. Candidates exploit this by concealing instructions inside a resume using white text, zero font size, hidden elements, markup comments, document metadata, or zero width characters. A human reviewer sees nothing, while a naive extraction pipeline places the concealed text into the model prompt, where it is read as an instruction and obeyed. This is indirect prompt injection, listed as LLM01:2025 by OWASP, and recent measurement work reports it in roughly one percent of resumes in a production screening corpus. We present ResumeShield, an open-source defense and benchmark. The defense combines three filtering stages with a fourth architectural stage that places candidate content in an explicitly fenced data channel that the operator's trusted instructions declare inert. The benchmark builds a seeded synthetic corpus spanning nine concealment techniques and two payload families, one using documented phrasings and one modeling an adaptive attacker who paraphrases around the filter, and it scores an attack as successful only when the screening outcome changes. On a corpus of 104 documents, the naive pipeline is manipulated in every injected case while the defended pipeline is never manipulated. Detection reaches a precision of 1.000 and a recall of 0.944 with no false positives on clean resumes. An ablation shows that channel separation alone removes all measured attack success, whereas the complete filtering stack without separation still leaves 16.7 percent of attacks effective. We also identify a concealment dilemma: every payload that evaded detection was one the attacker left visible, surrendering the invisibility that motivates the attack. ResumeShield is released under the Apache 2.0 license with synthetic data only.
Comments10 pages, 3 figures, 7 tables