分层网页证据投毒下的深度搜索智能体评估
Evaluating Deep-Search Agents under Hierarchical Web Evidence Poisoning
- Ant Group(蚂蚁集团)
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出HAE-GEO基准,通过三级网页证据投毒评估深度搜索智能体从暴露到恢复的完整轨迹,发现佐证陷阱削弱证据识别,防御提示难以将验证转化为恢复。
AI中文摘要:
搜索增强的LLM智能体越来越多地被用于消费者决策,这使得它们容易受到生成式引擎优化(GEO)投毒的攻击。现有基准大多衡量被操纵的内容是否被检索或认可,但并不追踪智能体是否验证可疑证据、修正已采纳的主张,或在产生最终推荐之前恢复。我们引入了HAE-GEO,一个在逐步更具说服力的网页投毒下,从暴露到恢复的完整轨迹追踪基准。智能体通过多轮搜索-抓取接口进行交互,跨越三个攻击级别(L1直接断言、L2上下文伪装和L3表面佐证),并由一个受控语料库支持,该语料库包含72,039个干净页面和每个级别770个投毒页面,涵盖8个产品类别和154个品牌。评估结合了确定性行为度量与六个语义评分维度。对10个智能体的评估揭示了三种反复出现的模式:在佐证陷阱下证据识别能力下降;智能体搜索提高了最终抵抗力,但没有提高证据识别或效用;防御性提示增加了验证,但很少将验证转化为恢复。
英文摘要:
Search-augmented LLM agents are increasingly used for consumer decisions, making them vulnerable to Generative Engine Optimization (GEO) poisoning. Existing benchmarks largely measure whether manipulated content is retrieved or endorsed, but do not track whether an agent verifies suspicious evidence, revises adopted claims, or recovers before producing its final recommendation. We introduce HAE-GEO, a benchmark that tracks the full trajectory from exposure to recovery under progressively more persuasive Web poisoning. Agents interact via a multi-turn Search-Scrape interface across three attack levels (L1 direct assertion, L2 contextual camouflage, and L3 apparent corroboration), supported by a controlled corpus of 72,039 clean pages and 770 poisoned pages per level spanning 8 product categories and 154 brands. Evaluation combines deterministic behavioral measures with six semantic rubric dimensions. Evaluating 10 agents, we find three recurring patterns: evidence recognition degrades under the corroboration trap; agentic search improves final resistance without improving evidence recognition or utility; and defense prompting increases verification, yet rarely converts verification into recovery.