arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MERA:面向大规模智能体系统的技能适配型模型演化与路由

MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale

Yuhang Yao, Zeyu Wang, Wanyi Chen, Tongyun Yang, Yuhang Han, Jie Xiao, Chengke Bao, Tianyi Zhao, Lynn Ai, Eric Yang, Tianyu Shi

arXiv 2608.10333首次发表:更新:

发表机构

Gradient; Soochow University; Carnegie Mellon University; Shanghai Jiao Tong University; University of California, Los Angeles(Gradient; 苏州大学; 卡内基梅隆大学; 上海交通大学; 加利福尼亚大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MERA以单模型调用为适应单元,通过多轮适配提升小模型能力,结合带验证器回退的成本校准路由,在HumanEval+MBPP、TAU-2等数据集上实现小模型性能提升与成本降低。

AI 中文摘要

大语言模型(LLM)智能体在单个任务中会执行异构的模型调用序列:部分调用需要细致推理,其余则是格式化或工具参数构建等结构化步骤。现有路由方法利用这种不对称性,将简单调用分配给低成本小模型,困难调用分配给大模型,这类策略可降低推理成本,但小模型能力保持不变,因此可实现的节省受限于小模型已能解决的任务量。MERA则以单个模型调用为适应单元来提升小模型自身能力:在每个循环中,MERA会重放小模型执行失败的调用,获取经执行验证的教师示范,将重复出现的过程提炼为迭代更新的《技能手册》(SkillBook),并通过监督学习和可选的GRPO微调小模型的LoRA适配器。路由作为部署的支撑机制:改进后的小模型由成本校准路由提供服务,该路由带有验证器支持的回退机制,且仅当联合重放能保持任务质量时,才会接纳候选的《技能手册》、适配器或路由。实验表明,四轮适应使Qwen2.5-Coder-1.5B在保留的HumanEval+MBPP数据集上的通过率从28.7%提升至49.7%;在验证器支持的回退机制下,部署策略以60.8%的全用Luna模型成本,保留了88.3%的通过率;在TAU-2数据集上,微调后的Qwen3.5-2B从14/35提升至18/35,性能与未适配的4B模型相当。这些结果表明,带验证器支持的多轮适应可提升小模型能力,而非仅围绕固定小模型进行路由。

英文摘要

LLM agents execute heterogeneous sequences of model calls within a single task: some invocations require careful reasoning, while others are structured steps such as formatting or tool-argument construction. Prior routing methods exploit this asymmetry by assigning easy invocations to a cheaper small model and difficult ones to a large model. Such policies reduce inference cost, but they leave the small model's capability unchanged, so attainable savings remain bounded by the work the student can already solve. MERA instead improves the small model itself, using a single model invocation as the unit of adaptation. In each cycle, MERA replays failed student invocations to obtain execution-verified teacher demonstrations, distills recurring procedures into an iteratively updated SkillBook, and fine-tunes a student LoRA adapter via supervised learning and optional GRPO. Routing serves as supporting machinery for deployment: the improved student is served behind a cost-calibrated router with verifier-backed fallback, and a candidate SkillBook, adapter, or router is admitted only when joint replay preserves task quality. Empirically, four-cycle adaptation raises Qwen2.5-Coder-1.5B from 28.7% to 49.7% pass on held-out HumanEval+MBPP. Under verifier-backed fallback, the deployed policy retains 88.3% pass at 60.8% of always-Luna cost. On TAU-2, a fine-tuned Qwen3.5-2B improves from 14/35 to 18/35 and matches an unadapted 4B model. These results indicate that verifier-backed multi-cycle adaptation can increase small-model capability, rather than only routing around a fixed student.

CommentsPreliminary version in CAIS RL-Eval

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑