AI 中文总结
研究跨语言提示注入攻击对基于大语言模型的相关性判断的影响,利用TREC数据集和开放权重模型在多语言环境下研究不同注入策略,发现多语言查询注入有效且能绕过现有防御,凸显当前防御差距,强调需更强大评估框架。
AI 中文摘要
大语言模型(LLMs)越来越多地被用作信息检索中相关性评估的自动评判工具,但其对对抗性操纵的鲁棒性仍未得到充分理解,尤其是在多语言环境中。在这项工作中,我们使用TREC深度学习数据集和两个开放权重模型,在既定的提示框架下,研究跨语言提示注入攻击对基于LLM的相关性判断的影响。我们在8种不同资源水平的语言中研究了基于指令和基于内容的注入策略。结果表明,基于多语言查询的注入在提高相关性分数的同时能有效规避现有提示注入防御。虽然现有防御机制可修改以减轻此类攻击,但这些注入可轻松适应以绕过它们。这些发现凸显了当前防御方法的关键差距,并表明语言泛化可成为攻击向量,强调需要为LLM评判系统建立更强大、主动的评估框架。
英文摘要
Large language models (LLMs) are increasingly being used as automated judges for relevance evaluation in information retrieval, yet their robustness to adversarial manipulation remains insufficiently understood, particularly in multilingual settings. In this work, we investigate the impact of cross-lingual prompt injection attacks on LLM-based relevance judgments using TREC Deep Learning collections and two open-weight models under established prompting frameworks. We examine both instruction-based and content-based injection strategies in 8 languages spanning different resource levels. Our results demonstrate that multilingual query-based injections are highly effective in inflating relevance scores while simultaneously evading existing prompt-injection defenses. We further found that, although existing defense mechanisms can be modified to mitigate such attacks, these injections can be easily adapted to bypass them. These findings highlight a critical gap in current defense approaches and demonstrate that language generalization can act as an attack vector, underscoring the need for more robust and proactive evaluation frameworks for LLM-as-a-judge systems.