arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26976cs.CLcs.IR

当学习到的上下文规划无法胜过强检索:规划、路由与重排在长上下文问答中的受控研究

When Learned Context Planning Fails to Beat Strong Retrieval: A Controlled Study of Planning, Routing, and Reranking for Long-Context QA

Yingrui Li, Han Chen

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过受控实验发现,学习到的上下文规划在长上下文多项选择问答中仅作为弱相关性信号,无法替代强检索,其性能在多数预算下不优于锚定混合检索或BM25。

中文摘要 AI 辅助

学习到的上下文规划在答案模型对证据原子进行推理之前选择它们。我们测试了这种学习到的选择是否能在强检索、路由、预算选择器和重排控制之后提升长上下文多项选择问答的性能。我们的主要诊断使用了全部503个LongBench-v2多项选择题,并采用Qwen2.5-7B-Instruct模型。规划器在来自140个训练问题和28个开发问题的基于结果选择的轨迹上进行SFT训练;由于503个问题的分析包含了这些问题,因此该分析部分上是传导性的。在18k字符预算下,锚定混合检索达到36.18%的准确率,BM25达到35.98%,而最佳的基于规划器直接引导的方法达到34.19%。在未触及的152个问题测试划分上,锚定混合仍然更高(42.11%对36.84%)。防泄漏路由器无法转化较大的预言机差距。在紧凑预算下,最佳规划器在6k时仅领先0.40个百分点,在9k时落后;规划器引导的重排在6k时有+1.79个百分点的估计,但配对区间跨越零,在9k时与控制组持平。打包顺序和分数平坦性分析未能识别出稳定的机制。在这种设置下,学习到的规划是一个弱的相关性信号,而非强检索的替代品。

英文摘要

Learned context planning selects evidence atoms before an answer model reasons over them. We test whether this learned selection improves long-context multiple-choice QA after strong retrieval, routing, budgeted-selector, and reranking controls. Our primary diagnostic uses all 503 LongBench-v2 MCQ questions with Qwen2.5-7B-Instruct. The planner is SFT-trained on outcome-selected traces from 140 training and 28 development questions; because the 503-question analysis includes those questions, it is partly transductive. At an 18k-character budget, anchored hybrid retrieval reaches 36.18% accuracy and BM25 reaches 35.98%, while the best direct planner-guided method reaches 34.19%. On the untouched 152-question test split, anchored hybrid remains higher (42.11% versus 36.84%). Leakage-safe routers cannot convert a large oracle gap. Under tight budgets, the best planner is ahead by only 0.40 points at 6k and loses at 9k; planner-guided reranking has a +1.79-point estimate at 6k with a paired interval crossing zero and ties the control at 9k. Packing-order and score-flatness analyses did not identify a stable mechanism. Under this setup, learned planning is a weak relevance signal rather than a replacement for strong retrieval.

补充信息

↑