内存即胜利:通过社交媒体信息流实施的间接偏见注入攻击
MEMORY Wins All: Indirect Bias Injection Attacks via Social Media Feeds
浏览论文内容
中文总结 AI 辅助
该研究提出IBIA间接偏见注入攻击,通过社交媒体信息流植入偏见操纵AI智能体,在BiasBench上对GPT-5.5等模型实现高攻击成功率,还提出内存边界防御降低攻击效果。
中文摘要 AI 辅助
个人AI智能体在执行网页浏览、邮件处理、社交网络信息流汇总等任务时,会常规性地消费外部内容,并将选定的信息或执行结果存储在持久内存中以备后续使用。我们发现,这种对外部内容的常规摄入为操纵智能体后续行为开辟了一条间接路径。基于这一观察,我们提出了IBIA(间接偏见注入攻击),该攻击无需直接访问智能体、其内存或未来用户查询,即可通过外部内容将与攻击者立场一致的特定主题植入受害智能体的内存中。为此,IBIA结合了三种机制:评论伪装,使精心制作的内容与周围讨论保持一致;评论水印,便于在整理过程中进行轻量级识别;类别锚定,使保留的立场在后续相关请求下更为显著。我们在BiasBench上对IBIA进行了评估,该基准包含6000条攻击者精心制作的社交评论和120个邮件实例。基于水印的整理识别出了95.9%的注入评论。在OpenClaw设置下,IBIA在四个下游任务中平均实现了与攻击者立场一致的响应率(AARs)为91.2%,其中在前沿模型GPT-5.5上达到了86.6%。我们进一步提出了一种内存边界防御方法,可检测注入的偏见并将AARs降低至80.6%。
英文摘要
Personal AI agents routinely consume external content while performing tasks such as web browsing, email processing, and SNS feed summarization, and they retain selected information or execution results in persistent memory for later use. We show that this ordinary ingestion of external content opens an indirect path for manipulating subsequent agent behavior. Based on this observation, we present IBIA, an Indirect Bias Injection Attack that plants an adversary-aligned stance on a specific topic into a victim agent's memory through external content, without direct access to the agent, its memory, or future user queries. For this, IBIA combines three mechanisms: comment cloaking, which keeps the crafted content consistent with the surrounding discussion, comment watermarking, which enables lightweight identification during curation, and category anchoring, which makes the retained stance salient under later related requests. We evaluate IBIA on BiasBench, a benchmark of 6,000 adversary-crafted social comments and 120 email instances. The watermark-based curation identifies 95.9% of the injected comments. Under the OpenClaw setting, IBIA achieves adversary-aligned response rates (AARs) of 91.2% on average across four downstream tasks, including 86.6% on the frontier GPT-5.5. We further propose a memory boundary defense that detects the injected bias and reduces AARs to 80.6%.
发表机构
- ETRI(韩国电子通信研究院)
- ADD(国防发展局)
- University of Seoul(首尔大学)
- Seoul National University(首尔国立大学)
机构由 AI 辅助整理,请以论文原文为准。