组装 CREW:用于自动相关工作生成的协作式多智能体强化学习框架
Assembling the CREW: A Collaborative Multi-agent Reinforcement Learning Framework for Automated Related Work Generation
查看机构详情
- The Chinese University of Hong Kong(香港中文大学)
- Hanoi University of Industry(河内工业大学)
- Hanoi University of Science and Technology(河内理工大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对多智能体LLM相关工作生成中静态协调的局限,提出CREW框架,让智能体通过IPPO策略自主选择操作动态协作,在标准基准上显著提升质量并降低令牌成本。
中文摘要 AI 辅助
自动相关工作生成(RWG)显著减少了撰写研究论文相关工作部分(RWS)所需的人力和时间。然而,先前利用多智能体大语言模型(LLM)的方法通常依赖预定义的工作流程,其中每个智能体负责整个过程中的特定步骤。这种僵化、静态的智能体间协调限制了综合复杂科学文献所需的适应性协作。为解决这一局限,我们提出了 CREW(用于相关工作生成的协作式强化学习),这是一个新颖框架,其中 LLM 智能体绕过启发式流水线,通过自主选择操作(如检索、传播、撰写和批判)来动态协调,并由通过独立近端策略优化(IPPO)优化的策略驱动。在标准 RWG 基准上的大量实验表明,我们的方法在显著降低令牌成本的同时,相较于强基线实现了大幅质量提升。代码可在以下 https URL 获取。
英文摘要
Automatic Related Work Generation (RWG) significantly reduces the human time and effort required to author the Related Work Section (RWS) of a research paper. However, prior methods leveraging multi-agent Large Language Models (LLMs) typically rely on a predefined workflow, where each agent is responsible for a specific step in the entire process. This rigid, static inter-agent coordination limits the adaptive collaboration required to synthesize complex scientific literature. To address this limitation, we propose CREW (Collaborative Reinforcement Learning for Related Work Generation), a novel framework where LLM agents bypass heuristic pipelines to dynamically coordinate by autonomously selecting actions, such as Retrieve, Disseminate, Compose, and Critique, driven by a policy optimized via Independent Proximal Policy Optimization (IPPO). Extensive experiments on a standard RWG benchmark demonstrate that our approach yields substantial quality improvements over strong existing baselines, while significantly reducing token costs. Code is available at https://github.com/YenPBao/CREW-Collaborative-MARL.git