arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Prompt2Skill:从自然语言指令进行无监督技能优化

Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructions

Bo Ni, Li Li, Ryan A. Rossi, Franck Dernoncourt, Tyler Derr

arXiv 2609.38593首次发表:更新:

发表机构

Vanderbilt University; University of Southern California; Adobe Systems(范德堡大学; 南加州大学; Adobe系统公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Prompt2Skill提出仅从自然语言任务描述构建技能的无监督框架,通过推导任务规范、合成数据集和闭环反思编辑,在四个领域上平均提升10.8个百分点,优于直接提示基线。

AI 中文摘要

技能是大型语言模型(LLMs)在推理时使用的外部构件,通过整合相关的程序性和领域知识来提升其在专业领域的性能。专家撰写的技能制作成本高昂,且生成的构件并未针对使用它的特定模型进行优化,因为模型的失败模式可能因版本、规模和训练而异。此外,新兴任务可能超出现有技能库的范围,从而需要在精选训练数据可用之前开发新技能。近期研究通过反思探索了自动化技能优化,但这些方法需要精选的、分布内的训练集,而用户可能并不总是拥有这样的训练集。为解决这些局限性,我们提出了Prompt2Skill,一个仅从自然语言任务描述构建技能的框架。该系统从提示中推导出任务规范,发现或合成数据集,并在反思性编辑的闭环中完善技能。在涵盖问答、阅读理解、电子表格操作和数学推理的四个领域中,Prompt2Skill始终优于直接提示基线,在开源模型和前沿模型上平均提升了10.8个百分点。

英文摘要

Skills are external artifacts that Large Language Models (LLMs) consume at inference time to improve their performance on specialized domains by incorporating relevant procedural and domain knowledge. Expert-authored skills are expensive to produce, and the resulting artifacts are not optimized for the specific model that consumes them, whose failure modes can vary with version, scale and training. In addition, emerging tasks may fall outside the scope of existing skill libraries, creating a need to develop new skills before curated training data become available. Recent works have explored automated skill optimization through reflection, but they require a curated, in-distribution training set, which users might not always have. To address these limitations, we present Prompt2Skill, a framework that builds skills from natural-language task description alone. From the prompt, the system derives a task specification, discovers or synthesizes datasets, and refines the skill in a closed loop of reflective editing. Across four domains spanning question answering, reading comprehension, spreadsheet manipulation, and mathematical reasoning, Prompt2Skill consistently outperforms the direct prompting baseline, achieving an average improvement of 10.8 across open-source and frontier models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑