arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

“我的名字有很多”:通过稀疏自编码器剖析助手及其角色设定

"Many Are My Names": The Anatomy of the Assistant and Its Personas via Sparse Autoencoders

Adelaide Danilov, Aria Nourbakhsh, Oleksandr Marchenko Breneur, Salima Lamsiyah

arXiv 2608.07852首次发表:更新:

发表机构

University of Luxembourg(卢森堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究借助稀疏自编码器剖析语言模型对说话者的内部表征,发现助手与角色扮演角色特征核心相关,故事角色则无,且可通过沉浸式模拟模式区分三者,助手有时会漂移至该模式。

AI 中文摘要

语言模型如何在内部表征说话者(包括助手、指定角色扮演角色或叙述故事中的角色)的问题仍未得到充分探索。本研究使用包含用户表达的情感文本及对应模型回复的数据集,通过过滤流程在不同深度选取轮次边界和代词标记位置提取的稀疏自编码器特征,将三种生成场景(助手、角色扮演、故事)进行分解。我们通过操控效应和激活分布表征每个留存特征,主要发现为:助手与角色扮演角色并非独立选项,角色设定保留与助手相关的特征核心,同时在各层逐步与助手区分,从操作机制转向行为和风格特征;而生成的故事角色缺乏与助手相关的核心。故事和角色扮演均可通过沉浸式模拟模式与助手区分,但助手有时甚至在默认设置下也会进入或逐渐漂移至该模式。

英文摘要

How a language model internally represents who is speaking, the Assistant, an assigned roleplay persona, or a narrated story character, remains underexplored. We study speaker representations using a dataset of user-expressed emotional text and corresponding model responses. We decompose three generation settings (Assistant, Roleplay, and Story) into sparse autoencoder features extracted at turn-boundary and pronoun-token positions and selected through a filtering pipeline for different depths. We characterize each surviving feature through its steering effects and activation distribution. Our main finding is that the Assistant and roleplay personas are not independent alternatives: personas retain the Assistant-associated feature core while progressively differentiating from it across layers, starting from operational machinery towards behavioral and stylistic features. Meanwhile, generated story characters lack the Assistant-associated core. Both Story and Roleplay can be distinguished from the Assistant with Immersive Simulation Mode. However, the Assistant can sometimes enter or slowly drift into it even in the default setting.

Comments38 pages, 9 tables, 4 figures, 2 listings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑