arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05383cs.AI

Sibyl:面向长时程任务的高效小-大模型协作框架

Sibyl: An Efficient Small-large Model Collaboration Framework for Long-horizon Tasks

Zhewei Fang, Yuxin Zhang, Zhenwei Shao, Mengze Li, Zheng Lin, Long Chen, Zhou Yu, Zhe Chen, Zhiwen Chen, Zhaode Wang, chengfei lv

首次发表
浏览论文内容

中文总结 AI 辅助

Sibyl提出三阶段训练框架,让0.6B小模型在长时程任务中通过选择性云端咨询和指导内化,以最少调用实现高成功率,显著超越现有基线。

中文摘要 AI 辅助

小语言模型(SLM)凭借低延迟、资源高效的推理能力,为端侧智能体提供了有前景的基础,但其有限的推理和规划能力制约了它们在需要与环境进行多步交互的长时程任务上的表现。SLM与云端大型模型之间的步骤级协作可以弥合这一差距,但识别需要云端协助的状态仍然具有挑战性:每次云端调用的贡献与后续动作纠缠在一起,只能从最终任务结果中评估。加剧这一挑战的是,SLM必须平衡两个相互竞争的目标:最大化任务成功率和最小化云端调用次数。为解决这一问题,我们提出了Sibyl,一种训练SLM智能体在步骤级别选择性咨询云端模型并将其指导内化用于后续决策的算法,从而以最小的云端依赖实现强大的任务性能。Sibyl遵循三阶段训练流程:(1)通过无咨询的自进化强化学习(RL)构建稳健的基础策略;(2)通过决定性分歧状态挖掘冷启动咨询行为;(3)通过咨询感知的强化学习联合优化咨询决策和指导内化。在ALFWorld和WebShop上的实验表明,仅使用0.6B参数模型的Sibyl在成功率上分别比最先进的基线(包括智能体训练和路由方法)高出95.2%和80.4%,同时每条轨迹平均仅需0.8次和3.9次云端调用。

英文摘要

Small language models (SLMs) offer a promising foundation for on-device agents through low-latency, resource-efficient inference, yet limited reasoning and planning capabilities constrain their performance on long-horizon tasks requiring multi-step interaction with the environment. Step-level collaboration between SLMs and larger cloud-hosted models can bridge this gap, but identifying states that warrant cloud assistance remains challenging: the contribution of each cloud call is entangled with subsequent actions and can be assessed only from the final task outcome. Compounding this challenge, the SLM must balance two competing objectives: maximizing task success and minimizing cloud calls. To address this, we propose Sibyl, an algorithm that trains SLM agents to selectively consult cloud models at the step level and internalize their guidance for subsequent decisions, achieving strong task performance with minimal cloud reliance. Sibyl follows a three-stage training pipeline that (1) builds a robust base policy through consultation-free self-evolving reinforcement learning (RL); (2) cold-starts consultation behavior via decisive-disagreement state mining; and (3) jointly optimizes consultation decisions and guidance internalization through consultation-aware RL. Experiments on ALFWorld and WebShop demonstrate that Sibyl, using only a 0.6B-parameter model, outperforms state-of-the-art baselines, including agent training and routing methods, by 95.2% and 80.4% in success rate while averaging only 0.8 and 3.9 cloud calls per trajectory, respectively.

发表机构

  • Hangzhou Dianzi University(杭州电子科技大学)
  • Fudan University(复旦大学)
  • The Hong Kong University of Science and Technology(香港科技大学)
  • University of Luxembourg(卢森堡大学)
  • Simon Fraser University(西蒙菲莎大学)
  • Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

↑