arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.31084cs.CL

首个 token 是线索:从 J-lens 表述多 token 概念

Verbalizing Multi-Token Concepts in LLMs

Xijie Gong, Zimeng Huang, Tonghan Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对 J-lens 无法表示多 token 概念的局限,提出利用 J-lens 的首个 token 线索结合冻结模型恢复多 token 概念及其向量的方法,在多个模型的多跳完形和因果概念交换任务上均优于 Template Lens。

中文摘要 AI 辅助

Jacobian Lens(J-lens)是近期用于解释大语言模型(LLMs)的工具,它将隐藏状态读取为词汇 token 的排序列表,导致多 token 概念缺乏自身的表示。原始 J-lens 研究通过 Template Lens(预计算固定短语词汇表的向量)和 Oracle Lens(微调组件以提出短语并重建短语向量)解决这一局限。本文探究是否可直接从 J-lens 和冻结模型中恢复多 token 概念及其向量。研究发现,多 token 概念的首个 token 与单 token 概念的可读性相当;在双 token 案例中,给定正确的首个 token 和源提示,冻结模型恢复第二个 token 的比例达 88.3%。研究表明,完整概念的向量可在单次前向传播中从后续隐藏状态恢复。因此,本文利用 J-lens 提出首个 token,让冻结模型补全候选概念,再恢复每个候选的向量并与完整词汇表一同评分。在 Gemma-3-12B-IT、Llama-3.1-8B 和 Qwen3-14B 上的 496 个多跳完形填空任务中,本文方法的平均 Rank@10 达 43.1%,而 Template Lens 为 27.6%;若无 J-lens 线索,性能降至 21.6%,表明首个 token 线索显著提升了读取效果。利用恢复向量进行因果概念交换的平均 succ@10 达 61.4%,而 Template Lens 在相同干预下为 26.2%。这些结果显示,首个 token 线索可引导多 token 概念恢复,后续隐藏状态则为读取和干预提供向量。

英文摘要

Lens methods inspect model computation by mapping intermediate activations to vocabulary tokens. Yet the concepts humans need to read out often span multiple tokens---entities, phrases, intermediate objects---making token-level readouts incomplete. Reliable multi-token readout with little model-specific preparation remains challenging. We introduce Concept Lens: token-level lens clues guide candidate concept search, then the model derives a representation for each candidate and scores it against the original activation. Across 2,400 multi-hop clozes on five LLMs (8B--70B), Concept Lens instantiated with J-lens and R-lens achieves average Rank@10 scores of 36.6\% and 54.5\%, respectively, compared with 21.7\% for Template Lens. Concept-swap interventions on derived concept representations shift model answers toward those associated with the replacement concepts. Further experiments show that Concept Lens can also reveal what a model recognizes along the way, beyond what appears in its final answer. Our code is available at https://github.com/XijieGo/c-lens

发表机构

  • College of AI, Tsinghua University(清华大学人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

↑