AI 中文总结
研究以搜索 API 为决策表面,用冻结的 GPT - 5.4 代理、两个工具及 100 个问题测试不同搜索提供商。发现虽答案准确性相近,但证据经济性不同,还引入新比率,指出提供商选择是检索预算和策略决策,非单纯召回决策。
AI 中文摘要
搜索 API 是许多代理的基本检索层且常用。传统搜索 API 提供网站内容预览。因全页检索消耗令牌多,代理检索架构常用渐进式披露。本文认为商业搜索 API 应视为决策表面。通过一个冻结的 GPT - 5.4 代理、两个工具及 100 个来自 SEALQA - HARD 的问题测试,不同搜索提供商(Brave、Tavily、Firecrawl)在答案准确性上相近,但证据经济性差异大。还引入表面矛盾与黄金 URL 比率。结果表明提供商选择是检索预算和策略决策,而非仅召回决策。
英文摘要
Search APIs expose ranked snippets, URLs, and metadata on which agents decide whether to answer, search again, or fetch pages. We evaluate these interfaces as decision surfaces on a fixed sample of 100 questions from the 254-question SealQA-Hard subset, using one frozen GPT-5.4 agent, a fixed orchestration harness, and a shared page-fetch backend across Brave, Tavily, and Firecrawl. A Kimi-K2.6 oracle labels visible URL-level evidence; a separate answer audit yields 25, 25, and 26 correct answers out of 100. These counts indicate similar observed accuracy but do not establish equivalence. Under the tested configurations, Brave exposes more pre-fetch support alongside a larger snippet surface; Tavily has a larger rank-1 share among trajectory-pooled supporting observations; and Firecrawl is associated with broader exploration. First-search query-level metrics distinguish support availability from ranking, while contradiction exposure complements contradiction-to-gold ratios. A retrospective oracle union covers 44/100 questions, versus 26/100 for the best individual provider: 18 percentage points of headroom. The observed evidence and action differences motivate evaluating search APIs jointly with agent policy and retrieval budget.