OPERA: A Reinforcement Learning--Enhanced Orchestrated Planner-Executor Architecture for Reasoning-Oriented Multi-Hop Retrieval
OPERA: 一种增强强化学习的协调规划-执行架构用于面向推理的多跳检索
机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) ; School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) ; School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院) ; Department of Electronic Engineering, Tsinghua University(清华大学电子工程系)
专题命中 长文档RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI
AI总结 OPERA通过协调规划-执行架构解决多跳检索中推理规划、检索和过滤的不足,采用MAPGRPO方法提升复杂任务性能。
Comments Accepted by AAAI 2026. Extended version