SkillSmith:通过自动技能构建与演进增强本地部署智能体
SkillSmith: Enhancing Locally Deployed Agents via Automatic Skill Construction and Evolution
浏览论文内容
中文总结 AI 辅助
SkillSmith是云端-本地智能体协作框架,通过自动构建和演进技能增强本地智能体,使搭载Qwen3.6-27B的本地智能体任务有效性接近云端智能体,在AppWorld上平均动作数从36.1降至9.9且可泛化至其他SLM。
中文摘要 AI 辅助
基于大语言模型(LLM)的智能体框架如今可作为个人助手完成多步骤任务。现有智能体框架如OpenClaw通常采用云端智能体部署模式,以闭源云端LLM作为骨干模型,这可能会泄露用户隐私信息并产生重复的LLM调用成本。本地智能体通过在用户控制的设备上部署前沿开源小型语言模型(SLM)解决了这些部署问题,但其任务有效性仍远落后于云端智能体。通过诊断分析,我们发现采用前沿SLM骨干模型的本地智能体有效性有限,主要源于骨干模型规模受限导致的环境知识缺失,包括环境规则和操作流程。为以非参数化、上下文高效且无需专家编写的方式提供此类知识,我们提出SkillSmith,这是一个云端-本地智能体协作框架,它将技能作为上下文高效的知识载体,通过云端智能体的任务探索自动构建技能,并利用本地智能体的执行反馈演进技能,以增强冻结的本地智能体。在日常智能体任务数据集AppWorld和WorkBench上的实验表明,自动生成的技能使搭载Qwen3.6-27B(SLM)的本地智能体达到了与搭载前沿LLM的云端智能体相当的任务有效性,优于最强的非参数化基线,在AppWorld-Normal上将每个任务的平均动作数从36.1减少到9.9,且无需重新运行技能构建即可泛化到其他SLM骨干模型。
英文摘要
LLM-based agent frameworks now act as personal assistants for multi-step tasks. Existing agent frameworks such as OpenClaw commonly follow the Cloud Agent depolyment mode using closed-source cloud LLMs as backbone model, which may expose private user information and incur repeated LLM-calling costs. Local Agents address these deployment concerns by depolying frontier open-source SLMs on user-controlled devices, but their task effectiveness still lags far behind Cloud Agents. Through diagnostic analysis, we reveal that the limited effectiveness of Local Agents with frontier SLM backbones mainly comes from missing environment knowledge caused by limited backbone model scale including environment rules and operation procedures. To supply such knowledge non-parametrically, context-efficiently, and without expert authoring, we present SkillSmith, a Cloud--Local Agent collaboration framework that uses Skill as a context-efficient knowledge carrier, automatic constructs Skill from Cloud Agent task exploration and evolves Skill using Local Agent execution feedback to enhance a frozen Local Agent. Experiments on daily agent task datasets AppWorld and WorkBench show that the automatically generated Skill enables the Local Agent with Qwen3.6-27B(SLM) to achieve task effectiveness comparable to Cloud Agents with frontier LLMs, outperform the strongest non-parametric baselines, reduce average actions per task from 36.1 to 9.9 on AppWorld-Normal, and generalize to other SLM backbone models without rerunning Skill construction.