arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29459cs.AI

面向潜在语言模型技能的引导与优化:一项实证研究

Toward Latent Language Model Skills Steering and Optimization: An Empirical Study

Xunyi Jiang, Junda Wu, Yuxin Xiong, Sheldon Yu, Tong Yu, David Arbour, Ritwik Sinha, Julian McAuley, Hongyi Wen

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过实证探究发现大型语言模型(LLM)的过程技能可表征为激活空间中的向量方向,且可通过向量操作实现技能引导与优化,为LLM过程技能的潜在空间操控提供了表征级视角。

中文摘要 AI 辅助

技能是大型语言模型(LLM)过程能力的有用抽象,体现模型执行结构化多步推理与程序的能力。现有方法通常将技能视为通过提示或程序指定的显式表层结构,未明确模型内部如何表征此类过程能力、能否在潜在空间中作为结构化对象被操控。本实证研究探究LLM过程技能是否可表征为激活空间中的方向,且这些方向上的向量空间操作能否表达技能级行为。研究发现,过程技能具有向量空间表征:单个技能方向可被激活以改变模型行为;独立提取的方向可组合形成更高级技能;对比方向可实现上下文条件下的算法个性化;技能方向上的优化轨迹呈非单调变化,中间状态常优于完全优化的解决方案。这些结果支持LLM过程技能的表征级视角:它们具有潜在向量空间组织,可通过内部干预直接操控。

英文摘要

Skills, as a useful abstraction for the procedural capabilities of large language models (LLMs), capture how models perform structured, multi-step reasoning and program execution. Existing approaches typically treat skills as explicit, surface-level constructs specified through prompts or programs, leaving open the question of how such procedural capabilities are represented inside the model and whether they can be manipulated as structured objects in latent space. In this empirical study, we investigate whether procedural LLM skills can be represented as directions in activation space and whether vector-space operations over these directions can express skill-level behaviors. We find that procedural skills admit a vector-space representation: individual skill directions can be activated to shift model behavior; independently extracted directions can compose to form higher-level skills. Contrastive directions yield context-conditioned algorithmic personalization and optimization trajectories over skill directions evolve non-monotonically, with intermediate states often surpassing fully optimized solutions. These results support a representation-level view of procedural LLM skills: they admit a latent vector-space organization that allows direct manipulation through internal interventions.

发表机构

  • New York University(纽约大学)
  • Adobe Research(奥多比研究院)
  • UC San Diego(加州大学圣迭戈分校)

机构由 AI 辅助整理,请以论文原文为准。

↑