超越提示或技能?基于归因的模块化语言程序优化
Beyond Prompt or Skill? Attribution-Guided Optimization of Modular LLM Programs
浏览论文内容
中文总结 AI 辅助
本文提出SPARO框架,通过归因引导的组件级优化,联合优化提示、技能与路由,在五个基准上超越现有基线,证明语言程序优化需同时关注知识发现与存储激活位置。
中文摘要 AI 辅助
大型语言模型能够解决日益多样化的推理任务,但其性能对任务提示、中间指令以及可复用问题解决知识的整合方式高度敏感。现有优化方法通常仅关注这一设计空间中的某一部分:它们要么优化整体提示,要么从模型轨迹中单独归纳和精炼技能。因此,这些方法缺乏在失败发生时决定应更新哪个组件的原则性机制,并且很少在统一框架中优化提示、技能和技能使用策略。我们提出SPARO(技能、提示与路由优化),这是一个联合优化任务指令、可复用技能块和路由规则的框架。它执行受控反事实评估,将示例的影响转化为提示、技能和路由组件上的概率责任分布,从该分布中采样一个组件,并应用相应的定向变异。这种设计将语言程序优化从全局提示重写转向结构化、可复用且选择性激活的任务知识。在五个基准测试和五个工作模型上,SPARO始终优于以提示为中心和以技能为中心的优化基线。这些结果表明,有效的语言程序优化不仅依赖于发现有用的任务知识,还取决于决定这些知识应存储在哪里以及何时应被激活。
英文摘要
Large language models can solve increasingly diverse reasoning tasks, yet their performance remains highly sensitive to task prompts, intermediate instructions, and the way reusable problem-solving knowledge is incorporated. Existing optimization methods usually focus on only one part of this design space: they either optimize a monolithic prompt, or separately induce and refine skills from model traces. As a result, they lack a principled mechanism for deciding which component should be updated when failures occur, and they rarely optimize prompts, skills, and skill-use policies in a unified framework. We propose SPARO (Skill, Prompt, And Routing Optimization), a framework that jointly optimizes task instructions, reusable skill blocks, and routing rules. It performs controlled counterfactual evaluations, converts examples' effects into a probabilistic responsibility distribution over prompt, skill, and routing components, samples one component from that distribution, and applies the corresponding targeted mutation. This design moves language-program optimization beyond global prompt rewriting toward structured, reusable, and selectively activated task knowledge. Across five benchmarks and five worker models, SPARO consistently outperforms both prompt-centered and skill-centered optimization baselines. These results suggest that effective language-program optimization depends not only on discovering useful task knowledge, but also on deciding where that knowledge should be stored and when it should be activated.
发表机构
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- Shenzhen Future Network of Intelligence Institute (FNii-Shenzhen)(深圳未来智能网络研究院(深圳FNii))
- Sun Yat-sen University(中山大学)
机构由 AI 辅助整理,请以论文原文为准。