自精化流水线中的非对称容量分配
Asymmetric Capacity Allocation in Self-Refinement Pipelines
浏览论文内容
中文总结 AI 辅助
本研究针对自精化流水线开展分阶段模型规模研究,发现生成、修订阶段需更大模型,批判阶段对规模不敏感,为多阶段语言模型系统的非对称容量分配提供了实用指导。
中文摘要 AI 辅助
自精化通常由生成、批判、修订三个阶段构成,是提升大语言模型(LLM)生成能力的广泛采用的范式,也是许多LLM智能体的核心机制。尽管这三个阶段涉及不同的认知需求,但现有多数方法将模型规模视为实现细节而非研究对象,可能导致资源浪费。鲜有研究系统考察模型规模对各阶段的影响,或有效的自精化是否需要生成、批判、修订阶段使用能力相当的模型。我们针对5个不同领域的基准测试,使用6种规模的Qwen3模型和4种规模的Gemma 3模型,开展了首个针对自精化流水线的分阶段模型规模研究。我们得出两个结论:其一,更大的生成器和修订器通常能提升流水线性能,而规模过小的修订器甚至会损害性能;其二,性能对批判模型的规模高度不敏感,不过即便使用小型批判模型,其效果也始终优于完全省略批判环节。我们的研究表明,模型容量不应在自精化流水线中均匀分配,不同阶段呈现出截然不同的规模缩放特性,为设计计算效率更高的多阶段语言模型系统提供了实用指导。
英文摘要
Self-refinement, typically structured as generation, critique, and revision, is a widely adopted paradigm for improving LLM generation and serves as a core mechanism in many LLM agents. While the three stages involve different cognitive demands, most existing approaches conveniently treat the model size as an implementation detail rather than a subject of study, which may lead to a waste of resources. Little work has systematically examined how model size affects each stage or whether effective self-refinement requires equally capable models for generation, critique, and revision. We present the first stage-wise model size study of the self-refinement pipeline on 5 benchmarks from different domains using 6 model sizes of Qwen3 and 4 model sizes of Gemma 3. We conclude that larger generators and refiners generally improve the pipeline, whereas an undersized refiner can even harm performance. Second, performance is highly insensitive to the size of the critic, although including even a small critic consistently outperforms omitting critique altogether. Our findings demonstrate that model capacity should not be allocated uniformly across self-refinement pipelines. Instead, different stages exhibit distinct size scaling characteristics, providing practical guidance for designing more computationally efficient multi-stage language model systems.
发表机构
- Drexel University(德雷塞尔大学)
- University of California Irvine(加州大学欧文分校)
机构由 AI 辅助整理,请以论文原文为准。