检索但未排序:结构检索中的表面形式偏差——从数学到智能体轨迹
Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to agent trajectories
浏览论文内容
中文总结 AI 辅助
该研究在数学和智能体轨迹领域发现嵌入检索存在表面形式偏差,LLM重排序器可部分缓解此偏差,而词汇重排序器的效果随领域变化,下游实验显示求解器准确率存在瓶颈。
中文摘要 AI 辅助
我们对嵌入检索进行评估,其中刻意将表面形式与含义分离:在一个协议下的两个不相关领域,检索共享底层结构但措辞不同的项目,分别是竞赛数学(MathNet-Retrieve;500个查询,117088项语料库)和具身智能体轨迹(源自ALFWorld;118个查询,336条轨迹)。在数学领域,失败是彻底的:在伪装程度最高的层级,两个生产级嵌入模型的严格Hit@1均为0.0%(自助法95%置信区间[0.0, 0.0]),而正确项目几乎总位于前10名;在95.2%至99.8%的错误案例中,获胜者比正确答案与查询的词汇相似性更高。在轨迹领域,表面变化是偶然的,当正确结果必须涉及不同对象时,相同模型的表现达到或接近超几何概率;当正确结果必须在对象和容器上都不同时,所有三个嵌入模型的表现均低于概率:检索锚定在字面标记而非任务结构上。词汇重排序器在数学领域有害,在轨迹领域有帮助(缩小26%至36%的差距,置信区间不包含0);其符号揭示基准的表面变化是对抗性的还是偶然的。大语言模型(LLM)重排序器在数学领域缩小5%至63%的差距,在轨迹领域缩小43%至76%的差距;方向在三位评判者间一致(所有21个单元格均为正),但效应量、层级分布和异常评判者随领域变化(配对差异处处不包含0)。数学领域的增益集中在知名竞赛(+19.8个百分点,置信区间[+6.7, +33.2],为6个单元格之一),因此部分恢复源于记忆。在配对下游实验(210个查询,评判者一致性为96%至99%)中,神谕检索与对抗性差检索无差异(McNemar检验p=0.678);求解器的69.5%零样本准确率在很大程度上是截断代理(对完整答案的准确率为97%至100%),无提升空间。
英文摘要
We evaluate embedding retrieval where surface form and meaning are pulled apart on purpose: retrieving items that share underlying structure but not wording, in two unrelated domains under one protocol, competition mathematics (MathNet-Retrieve; 500 queries, 117,088-item corpus) and embodied-agent trajectories (ALFWorld-derived; 118 queries, 336 trajectories). In mathematics the failure is complete: strict Hit@1 at the heaviest disguise tier is 0.0% for both production embedders (bootstrap 95% CI [0.0, 0.0]) while the correct item sits in the top 10 nearly always, and in 95.2 to 99.8% of misses the winner is more lexically similar to the query than the correct answer. In trajectories, where surface variation is incidental, the same models land at or near hypergeometric chance when gold must involve a different object, and below chance for all three embedders once gold must differ in object and receptacle: retrieval anchors on literal tokens, not task structure. A lexical reranker control hurts in mathematics and helps in trajectories (closing 26 to 36% of the gap, CIs excluding zero); its sign reveals whether a benchmark's surface variation is adversarial or incidental. An LLM reranker recovers 5 to 63% of the gap in mathematics and 43 to 76% in trajectories; direction replicates across three judges (all 21 cells positive), but effect sizes, tier profiles, and the outlier judge change with domain (paired differences excluding zero everywhere). Mathematics gains concentrate on well-known competitions (+19.8 points, CI [+6.7, +33.2], one of six cells), so part of the recovery is memorization. In a paired downstream experiment (210 queries, graders at 96 to 99% agreement), oracle retrieval was indistinguishable from adversarially bad retrieval (McNemar p = 0.678); the solver's 69.5% zero-shot accuracy is largely a truncation proxy (97 to 100% on finished answers), leaving no headroom.
发表机构
- MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)
- Mantis(曼蒂斯(机构名))
机构由 AI 辅助整理,请以论文原文为准。