AI供应链应用中的联合中毒攻击
Conjunctive Poisoning in AI Supply-Chain Applications
浏览论文内容
中文总结 AI 辅助
本研究提出AI供应链中的联合中毒攻击,通过包装器与元数据的交互改变模型行为,提出TIF-BAH防御,揭示AI部署中存在未被现有防御覆盖的部署时行为风险。
中文摘要 AI 辅助
大型语言模型(LLM)和视觉语言模型(VLM)越来越多地通过推理流水线部署,这些流水线包含提示包装器(如模板和后处理脚本)以及配置元数据(如JSON/YAML文件),二者共同塑造模型输出。尽管模型权重和二进制文件会被常规验证,但这些文本部署工件仍保护薄弱,却直接影响运行时行为。我们证明,恶意开发者可将看似良性的包装器与精心构造的元数据配对,无需修改模型权重、训练数据或推理后端即可确定性地改变生成后行为。我们通过受控的联合门实现研究该行为,其激活取决于嵌入的包装器标记和密码学绑定的元数据。我们在15个开源和闭源LLM/VLM部署上评估该攻击,并评估静态元数据检查、包装器扫描器、PromptShield和基于SigStore的工件签名等提示及系统级防御。为缓解此风险,我们提出TIF-BAH,一种轻量级中间件防御,可在推理期间验证包装器完整性并记录行为证明。我们的结果表明,包装器-元数据交互构成现代AI部署中保护不足的执行层,暴露出模型权重或提示级防御未覆盖的部署时行为风险。代码可在该https URL获取。
英文摘要
Large Language and Vision-Language Models are increasingly deployed through inference pipelines that include prompt wrappers (e.g., templates and post-processing scripts) and configuration metadata (e.g., JSON/YAML files) that together shape model outputs. While model weights and binaries are routinely verified, these textual deployment artifacts remain weakly protected despite directly influencing runtime behavior. We show that a malicious developer can pair a benign-looking wrapper with crafted metadata to deterministically alter post-generation behavior without modifying model weights, training data, or inference backend. We study this behavior through a controlled conjunctive-gate implementation, where activation depends on both an embedded wrapper marker and cryptographically bound metadata. We evaluate the attack across fifteen open- and closed-source LLM/VLM deployments, and assess prompt and system level defenses including static metadata inspection, wrapper scanners, PromptShield, and SigStore-based artifact signing. To mitigate this risk, we introduce TIF-BAH, a lightweight middleware defense that verifies wrapper integrity and records behavioral attestations during inference. Our results reveal that wrapper-metadata interactions form an under-protected execution layer in modern AI deployments, exposing a deployment-time behavioral risk that is not captured by model-weight or prompt-level defenses. Code is available at https://github.com/N-H-Arif/llm_temp.