发表机构
University of Massachusetts Amherst; Tel Aviv University(马萨诸塞大学阿默斯特分校; 特拉维夫大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TellTail通过查询指纹识别黑盒检索系统背后的嵌入模型,利用检索结果差异和查询迁移性,在多种访问级别下高精度识别,揭示知识产权泄露风险。
AI 中文摘要
密集嵌入模型是现代文本检索的核心,支撑着从网络搜索到检索增强生成(RAG)等系统。然而,检索器通常部署在不透明的系统中,仅暴露排名结果、引用来源或生成的答案。这种不透明性阻碍了用户验证提供商提供的是哪个检索器,并可能造成对检索攻击的虚假安全感。我们证明,攻击者仅通过查询就能推断出检索器的身份。我们提出了TellTail,一种指纹识别攻击,用于识别黑盒系统背后的检索器,覆盖一系列访问级别。尽管检索器共享许多属性,但它们可以被引导产生不同的结果。当检索到的段落被暴露时,TellTail比较检索重叠以推断底层模型。当仅暴露最终生成的响应时,TellTail优化特定于模型的查询,这些查询在目标检索器上诱导出选定的检索行为,但在其他检索器上迁移性差。有趣的是,限制攻击在其他地方迁移性的因素恰恰使它们对指纹识别有用。我们使用53个检索器在三种日益严格的设置下评估TellTail。TellTail从完整检索排名中完美识别部署的检索器;在无序前3结果中,94.3%的尝试成功;仅从语言模型生成的答案中,92.5%的情况成功。此外,指纹识别成功对抗了一个流行的RAG系统(OpenWebUI)。总体而言,TellTail表明基于检索的系统泄露了其嵌入模型的身份——暴露知识产权,并在针对特定模型的攻击(如语料库投毒)之前实现侦察。
英文摘要
Dense embedding models are core to modern text retrieval, enabling systems ranging from web search to retrieval-augmented generation (RAG). Yet, retrievers are usually deployed within opaque systems, exposing only ranked results, cited sources, or generated answers. This opacity prevents users from verifying which retrievers providers serve and may create a false sense of robustness against retrieval attacks. We show that attackers can infer retrievers' identity only through queries. We introduce TellTail, a fingerprinting attack for identifying retrievers behind black-box systems across a spectrum of access levels. Despite sharing many properties, retrievers can be steered to emit distinct results. When retrieved passages are exposed, TellTail compares retrieval overlap to deduce the underlying model. When only the final generated response is exposed, TellTail optimizes model-specific queries that induce a chosen retrieval behavior on the target retriever but transfer poorly to others. Interestingly, the poor transferability that limits attacks elsewhere is exactly what makes them useful for fingerprinting. We evaluate TellTail using 53 retrievers under three increasingly restrictive settings. TellTail perfectly identifies the deployed retriever from full retrieval rankings; in 94.3% of attempts with unordered top-3 results; and in 92.5% of cases from language-model-generated answers alone. Furthermore, fingerprinting succeeds against a popular RAG system (OpenWebUI). Overall, TellTail shows that retrieval-based systems leak the identity of their embedding model---exposing intellectual property and enabling reconnaissance ahead of model-specific attacks such as corpus poisoning.