接下来应该问什么?部分证据下的检索感知问题学习
What Should We Ask Next? Retrieval-Aware Question Learning for Interactive ReID
AI总结:
针对部分证据下的交互式检索,提出检索感知在线强化学习框架RAVEL,通过排名反馈优化问题策略,在行人重识别任务中逐步提升检索性能。
AI中文摘要:
部分证据下的交互式检索是一个序贯信息获取问题:智能体必须决定哪个问题能为下一次检索更新创造最有用的证据。现有系统通过模仿候选问答对的离线排序来训练这一决策,尽管问题的价值取决于其引发的回答及其对检索的下游影响。我们确定了候选区分度和感知有用性为该目标提供了弱监督,然后引入了RAVEL,一种用于交互式行人重识别的检索感知在线强化学习框架。RAVEL从监督式问题生成初始化,直接观察当前的Top-4候选,并通过来自完整问答-检索循环的排名反馈优化问题策略。在Interactive-PEDES上的实验表明,RAVEL在五轮交互中提供了逐步增强的检索性能。进一步分析显示,RAVEL将提问预算重新分配到局部开放属性上,这些属性提供了更有用的检索证据,并在最初困难的查询上产生了最大的收益。
英文摘要:
Interactive retrieval with partial evidence constitutes a sequential information-acquisition problem: an agent must choose questions that acquire useful evidence for the next retrieval update. Existing systems train this decision by imitating an offline ordering of candidate QA pairs. However, a question's value depends on the response it elicits and its downstream effect on retrieval. We establish that candidate discriminativeness and perceived usefulness provide weak supervision for this objective, then introduce RAVEL, a retrieval-aware online reinforcement learning framework for interactive person re-identification. RAVEL initializes from supervised question generation, observes the current Top-4 candidates directly, and optimizes the question policy with rank feedback from the full question-answer-retrieval loop. Experiments on Interactive-PEDES show that RAVEL delivers progressively stronger retrieval performance across five interaction rounds. Further analysis shows that RAVEL allocates more of its interaction budget to localized open-ended prompts targeting specific attributes.