arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CPInj:揭示文本协作式提示优化中的提示注入风险

CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization

Xinting Liao, Behnoosh Zamanlooy, Masoumeh Shafieinejad, David B. Emerson, Ruinan Jin, Deval Pandya, Xiaoxiao Li

arXiv 2607.18622首次发表:更新:

发表机构

University of British Columbia; Vector Institute; McMaster University(不列颠哥伦比亚大学; 向量研究所; 麦斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究文本协作式提示优化(TCPO)中提示注入风险,提出协作式提示注入攻击CPInj,用恶意指令污染全局提示,降低下游任务性能,现有防御方法无效;还提出防御导向聚合方法APAgg,实验表明TCPO存在关键漏洞,攻击有待更强大防御。

AI 中文摘要

文本协作式提示优化(TCPO)通过让多个客户端在本地保留数据的同时共同改进大语言模型(LLM)的提示,将Textgrad扩展到分散式设置。其对自由形式文本更新和聚合的依赖引入了新的、未充分探索的攻击面,恶意指令可注入本地提示并通过服务器端提示聚合传播。不同于传统提示注入攻击,CPInj针对TCPO中的协作优化循环。我们提出CPInj,一种协作式提示注入攻击,它用恶意指令污染聚合的全局提示,降低下游任务性能,抵抗良性客户端的提示优化净化,并逃避服务器上基于高级检测的防御。当前防御方法对CPInj无效,我们进一步提出防御导向的聚合方法APAgg,可净化恶意指令并部分恢复TCPO效用。我们在三个LLM系列和数学、逻辑、医学的五个推理任务上进行了广泛实验。结果表明我们提出的攻击揭示了TCPO中的关键漏洞,尽管采取了缓解措施,但攻击仍然非常有效且远未完全解决,需要为TCPO提供更强大的防御。

英文摘要

Textual Collaborative Prompt Optimization (TCPO) extends TextGrad (Yuksekgonul et al., 2025) to a decentralized setting by allowing multiple clients to jointly improve prompts for large language models (LLMs) while keeping their data locally. Its reliance on free-form textual updating and aggregation introduces a new and largely unexplored attack surface, i.e., malicious instructions can be injected into local prompts and propagated through server-side prompt aggregation. Unlike conventional prompt injection attacks, attacking TCPO targets the collaborative optimization loop in TCPO. This setting is more challenging because malicious instructions must survive aggregation, persist through subsequent benign prompt optimization, and evade server-side defenses. To expose this risk, we propose Collaborative Prompt Injection (CPInj) attack that contaminates the aggregated global prompt with malicious instructions, degrades downstream task performance, resists purification by prompt optimization on benign clients, and evades advanced detection-based defenses on the server. We find that current defense methods are ineffective against CPInj. We further propose Anchored Purification Aggregation (APAgg), a defense-oriented aggregation that purifies malicious instructions without severely degrading TCPO utility. We conduct extensive experiments across three LLM families and five reasoning tasks in math, logic, and medicine, and demonstrate that our proposed attack reveals a critical vulnerability in TCPO. Although we take a first step toward mitigation, the attack remains highly effective and far from fully resolved, calling for more robust defense for TCPO.

CommentsAccepted by COLM 2026 and AI4GOOD workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑