arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06879cs.CLcs.AI

AutoLexSteer:用于词汇激活引导的自动对比构建

AutoLexSteer: Automatic Contrast Construction for Lexical Activation Steering

发表机构墨尔本大学计算与信息系统学院
查看机构详情
  • School of Computing and Information Systems, The University of Melbourne(墨尔本大学计算与信息系统学院)

机构由 AI 辅助整理,请以论文原文为准。

Shuhe Wang, Lachlan Cowley, Eduard Hovy, Jey Han Lau

首次发表
浏览论文内容

中文总结 AI 辅助

AutoLexSteer提出首个全自动构建词汇引导向量的方法,利用WordNet词族实现精确的词义级引导,并能调节谄媚等行为。

中文摘要 AI 辅助

引导向量已迅速成为一种流行且有效的方法,用于以非常具体的方式引导大型语言模型的输出。但由于嵌入的不透明性,构建准确的引导向量是一个困难的手动过程。我们引入了Hangman,一种基于词义运作的新型引导向量,以及AutoLexSteer,第一个完全自动化的引导向量构建流程。AutoLexSteer利用从WordNet中提取的紧密相关词族来指定要避免的引导源和期望的引导目标。这些引导向量非常精确,可用于在单词和词义集(含义)层面进行引导,并能够引导某些大型语言模型行为,如谄媚。数据集和代码可在以下网址找到:此https URL。

英文摘要

Steering vectors have rapidly emerged as a popular and effective method for guiding the output of LLMs in very specific ways. But constructing accurate steering vectors is a difficult manual process due to the opacity of embeddings. We introduce Hangman, a novel type of steering vector that operates using word senses, as well as AutoLexSteer, the first fully automated process for building steering vectors. AutoLexSteer employs families of closely-related words extracted from WordNet to specify both the steering source to be avoided and the desired steering target. The steering vectors are quite precise, can be used to steer at the level of words and sets of word senses (meanings), and are able to steer certain LLM behaviors like sycophancy. The dataset and code can be found at https://github.com/ShuheWang1998/autolexsteer.

↑