AI 中文总结
研究PRF失败问题,提出两阶段审计然后自动化框架,第一阶段通过对用户审计揭示PRF利弊,第二阶段利用基于语言模型的重排器作预测器,解释PRF危害及决策过程,使检索组件可审计且以用户为基础。
AI 中文摘要
伪相关反馈(PRF)平均提高检索效果,但因查询漂移损害了相当一部分查询,而聚合离线指标掩盖了这种不对称性。现有选择性PRF方法通常依赖相同排名统计的查询性能预测方法,继承而非解决了这种不透明性。我们提出两阶段审计然后自动化框架。第一阶段,对43个TREC 2019深度学习查询的108个用户进行参与式审计,发现仅20.9%的查询受益于PRF,25.6%用户体验下降,避免损害的价值几乎是利用成功扩展的两倍。第二阶段,将基于语言模型的重排器重新用作系统偏好预测器,基于可检查文档证据自动复制这些用户得出的标签。两个阶段共同解释了PRF对哪些查询有害、为何做出选择性PRF决策以及如何大规模检查该决策,将不透明的检索组件转变为可审计、以用户为基础的组件。
英文摘要
Pseudo-Relevance Feedback (PRF) improves retrieval effectiveness on average, but harms a substantial fraction of queries through query drift, an asymmetry hidden by aggregate offline metrics. Existing Selective PRF (sPRF) approaches typically rely on Query Performance Prediction (QPP) methods derived from the same ranking statistics, and therefore inherit, rather than resolve, this opacity. We argue that this is a core explainability problem in IR, and propose a two-stage audit-then-automate framework. In Stage 1, a participatory audit with 108 users across 43 TREC Deep Learning 2019 queries shows that only 20.9% of queries benefit from PRF, while 25.6% suffer a degraded user experience, and that avoiding harm is nearly twice as valuable as exploiting successful expansion. In Stage 2, we repurpose LLM-based rerankers as system preference predictors that replicate these user-derived labels automatically, grounded in inspectable document evidence. Together, the two stages explain which queries PRF harms, why an sPRF decision is made, and how the decision can be inspected at scale, turning an opaque retrieval component into an auditable, user-grounded one.
CommentsAccepted at WExIR @ SIGIR 2026, the 2nd Workshop on Explainability in Information Retrieval, Melbourne (Naarm), Australia, 24 July 2026. 3 pages, 1 figure. Extended abstract building on the SIGIR 2026 short paper "Auditing Query Drift: Do Users Actually Benefit from Pseudo-Relevance Feedback?" (doi:10.1145/3805712.3809916)