arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05475cs.LG

KV-Skill:在模型的原生语言中锻造专业技能

KV-Skill: Forging Expertise in the Model's Native Language

Zhaowei Han, Xiang Zhang, Bing Han, Kai Liu, Danqi Hu, Jie Liu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出KV-Skill,通过外部因子化算子存储任务知识,在多基准测试中提升了模型在LiveMath等任务上的准确率,且可独立加载多个技能无明显遗忘。

中文摘要 AI 辅助

任务知识通常存储为提示词中的文本,或作为模型权重的更新。文本具有模块化,但每次使用都必须被解释;而权重适配会使生成的能力难以独立加载、移除或共享。我们引入KV-Skill,这是一个外部因子化算子的设计空间,冻结的语言模型通过轻量级接口读取这些算子。KV-Skill支持两条互补路径:注册(Registration)将已编写的文本技能转换为文本衍生算子,并训练一个与主干模型相关的共享接口;奖励学习(Reward learning)则可直接从任务结果中开发紧凑的隐空间算子,无论是否有已编写的技能。两条路径均不会向提示词添加位置。在来自三个模型家族的四个主干模型和十个基准测试中,将文本转换为KV-Skill始终能使相同的过程知识更有效。在Qwen3.5-4B LiveMath上,注册达到77.2的准确率,而源文本技能为23.4,SkillOpt为52.0,SoftSkill为64.5。在匹配的奖励训练和参数预算下,与软前缀、前缀调优和LoRA相比,KV-Skill在八个匹配设置中的七个取得最佳结果。事后排名分析进一步表明,文本衍生算子在每个注入层使用一个任务对齐方向时,几乎保留了全部优势,而匹配的随机方向则无法做到。最后,一个共享接口可保留三个独立可加载的KV-Skill,且无明显遗忘。这些结果表明,任务知识可以从文本或经验中获取,压缩为外部算子,并与主干模型分开部署。代码可在以下网址获取:this https URL

英文摘要

Task knowledge is commonly stored either as text in the prompt or as an update to model weights. Text is modular but must be interpreted on every use, while weight adaptation makes the resulting capability difficult to load, remove, or share independently. We introduce KV-Skill, a design space of external factorized operators that a frozen language model reads through a lightweight interface. KV-Skill supports two complementary paths. Registration converts an authored text skill into a text-derived operator and trains a shared per-backbone interface. Reward learning develops a compact latent operator directly from task outcomes, with or without an authored skill. Neither path adds positions to the prompt. Across ten benchmarks and four backbones from three model families, converting text to a KV-Skill consistently makes the same procedural knowledge more effective. On Qwen3.5-4B LiveMath, registration reaches 77.2 accuracy, compared with 23.4 for the source text skill, 52.0 for SkillOpt, and 64.5 for SoftSkill. Under matched reward training and parameter budgets, KV-Skill gives the best result in seven of eight matched settings against soft prefixes, prefix tuning, and LoRA. A post-hoc rank analysis further shows that text-derived operators retain nearly all of their benefit with one task-aligned direction per injection layer, while matched random directions fail. Finally, one shared interface retains three independently loadable KV-Skills without measurable forgetting. These results show that task knowledge can be acquired from text or experience, compressed into an external operator, and deployed separately from the backbone. Code is available at: https://github.com/shawnzhg/KV-Skill

发表机构

  • University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑