arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

忠实激活言语化:减少大语言模型表示解释中的幻觉

Faithful Activation Verbalization: Reducing Hallucinations in LLM Representation Interpretation

Haiyan Zhao, Zirui Hei, Wei Shi, Huiqi Deng, Na Zou, Mengnan Du

arXiv 2609.34033首次发表:更新:

发表机构

New Jersey Institute of Technology; Shanghai AI Laboratory; Xi’an JiaoTong University; Chinese Univerity of Hongkong, Shenzhen(新泽西理工学院; 上海人工智能实验室; 西安交通大学; 香港中文大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AVPO通过两阶段框架(重建源文本并用冻结问答模型评估)结合直接偏好优化,显著减少激活言语化中的幻觉,提升信息恢复并超越基线。

AI 中文摘要

激活言语化方法(如Activation Oracle和Natural Language Autoencoders)将大语言模型的隐藏表示解码为人类可读的自然语言。然而,现有方法可能产生不完整或幻觉性的描述,使得其激活言语化在实践中难以信任和可靠使用。为此,我们提出了AVPO,一个两阶段框架,首先从隐藏激活中重建源文本,然后用一个独立的冻结问答模型评估生成的文本,从而产生一个明确且可检查的中间读出。我们进一步使用直接偏好优化(DPO)优化逆变器,利用同时捕获语义可恢复性和词汇保真度的奖励。在六个文本族中,AVPO在要点级和细节级信息恢复上分别比最强基线提高了最多17.1和9.3个百分点。关键在于,这些提升来自偏好优化而非仅对选定重建进行微调,使得紧凑的跨模型逆变器能够超越供体匹配的问题条件言语化器,同时提高语义可恢复性和词汇保真度。此外,分布外案例研究表明,AVPO能更好地恢复高层语义,同时减少编造的细节。

英文摘要

Activation verbalization methods such as Activation Oracle and Natural Language Autoencoders decode hidden representations of large language models into human-readable natural language. However, existing methods can produce incomplete or hallucinated descriptions, making their activation verbalizations difficult to trust and use reliably in practice. To this end, we introduce AVPO, a two-stage framework that first reconstructs source text from a hidden activation and then evaluates the resulting text with a separate frozen question-answering model, yielding an explicit and inspectable intermediate readout. We further optimize the inverter with direct preference optimization (DPO), using rewards that capture both semantic recoverability and lexical fidelity. Across six text families, AVPO improves gist- and detail-level information recovery over the strongest baseline by up to 17.1 and 9.3 percentage points, respectively. Crucially, the gains arise from preference optimization rather than fine-tuning on selected reconstructions alone, enabling compact cross-model inverters to surpass donor-matched question-conditioned verbalizers while improving both semantic recoverability and lexical fidelity. Moreover, out-of-distribution case study shows that AVPO better recovers high-level semantics while fabricating fewer details.

Comments34 pages, 13 figures, 13 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑