arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29215cs.CL

面向特定群体解释生成的大语言模型基于属性的激活导向

Attribute-Based Activation Steering of LLMs for Group-Specific Explanation Generation

Leandra Fichtel, Janek Prange, Henning Wachsmuth

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对特定群体解释生成需求,提出基于属性的LLMs激活导向方法,经实验验证其相比基线方法能更好适配目标群体并保持事实性。

中文摘要 AI 辅助

为让人们有效理解新主题,解释需适配其背景与能力,仅靠提示工程不足以生成此类解释,且缺乏其他计算方法,因此本文探究能否引导大语言模型(LLMs)生成适配特定群体的解释。为此,本文提出一种方法:先识别特定目标群体在解释风格与知识方面的群体专属属性,基于激活工程计算基于属性的导向向量,在推理阶段将其添加至LLMs的内部激活中,以实现细粒度导向。实验中,本文从生成解释的针对性与事实性评估该方法的导向效果,还邀请不同目标群体的人类专家开展研究以评估解释。与提示工程及现有最优导向基线方法相比,本文方法能显著更好地将解释适配至目标群体,同时在很大程度上保持事实性。

英文摘要

To effectively enable people to understand new topics, explanations should be tailored to their backgrounds and abilities. Prompting alone has been shown to be insufficient for creating such explanations, and other computational methods are missing so far. Therefore, this paper investigates whether LLMs can be steered to generate explanations that are tailored to a specific group of people. To this end, we propose an approach that first identifies group-specific attributes in terms of explanatory style and knowledge of a specific target group. Building on activation engineering, it then computes attribute-based steering vectors and adds them to the internal activations of an LLM during inference to enable a fine-grained steering. In our experiments, we assess the steering effectiveness in terms of specificity and factuality of the generated explanations. Additionally, we evaluate the explanations in a study with human experts from different target groups. Compared to prompting and state-of-the-art steering baselines, our approach tailors the explanations significantly better to the target group while maintaining the best specificity-factuality balance.

发表机构

  • Leibniz University Hannover(汉诺威莱布尼茨大学)
  • L3S Research Center(L3S研究中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑