发表机构
Meituan(美团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出 LongCat-DeepResearch,一种结合增强 LongCat 模型与多智能体工作流的深度研究系统,通过分离全局规划与局部调查、章节级修订,在多个基准上取得领先成绩,并支持模型训练。
AI 中文摘要
我们提出了 LongCat-DeepResearch,一个深度研究系统,它将增强的 LongCat 模型与多智能体工作流相结合,用于生成全面、有证据支撑的报告。该工作流将全局规划与详细调查分离,并在章节层面协调修订。多个规划智能体首先探索外部来源并完善一个可操作的研究计划,称为 ResearchSpec。随后,研究智能体并行调查并起草其分配的章节,随着分析的发展,在独立的上下文中收集额外证据。章节汇总后,全局审查指导有针对性的局部修订,减少对重复全文重写的依赖。该工作流还支持为 LongCat 通用模型的中期训练和后期训练构建研究任务和轨迹。LongCat-DeepResearch 在 DeepResearchBench 上达到 55.25,在 DeepResearchBench II 上达到 51.35,在 ResearchRubrics 上达到 79.83。在一个内部基准上,它得分 76.04,在四个比较系统中排名第二。开发集分析显示,结合规划视角有益,而进一步细化规划则效果不一。额外的编辑提高了两个基准上的平均自动可读性偏好,但各自趋势不同。
英文摘要
We present LongCat-DeepResearch, a deep research system that combines an enhanced LongCat model with a multi-agent workflow for producing comprehensive, evidence-grounded reports. The workflow separates global planning from detailed investigation and coordinates revision at the section level. Multiple planning agents first explore external sources and refine an actionable research plan, termed ResearchSpec. Research agents then investigate and draft their assigned sections in parallel, gathering additional evidence in separate contexts as their analyses develop. Once the sections are assembled, global review guides targeted local revisions, reducing reliance on repeated full-report rewriting. This workflow also supports the construction of research tasks and trajectories for the mid-training and post-training of LongCat's general-purpose models. LongCat-DeepResearch achieves 55.25 on DeepResearchBench, 51.35 on DeepResearchBench II, and 79.83 on ResearchRubrics. On an in-house benchmark, it scores 76.04, ranking second among four compared systems. Development-set analyses show benefits from combining planning perspectives, while further planning refinement has mixed effects. Additional editing improves average automatic readability preference across two benchmarks, with different trends on each.
Comments23 pages, 5 figures