arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LLM中的病感失认:探究量化计算底层的自我意识

Anosognosia in LLMs: Probing Self-Awareness of Quantized Computational Substrate

Yoshihiro Izawa, Gouki Minegishi, Yoko Yamakata

arXiv 2610.06174首次发表:更新:

发表机构

The University of Tokyo(东京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究受病感失认启发,探究LLM能否识别自身量化引起的计算底层退化,发现内部表示含方法特定指纹,共享LoRA可读取但泛化有限,提示内部表示是实现自我意识的关键途径。

AI 中文摘要

大型语言模型(LLM)能否识别自身计算底层(computational substrate)的退化?受病感失认(anosognosia)的启发——这是一种患者无法识别自身能力受损的神经系统疾病——我们研究LLM能否识别由量化(quantization)引起的计算底层退化。我们首先表明,现有模型无法自我报告其量化状态,即使提供其自身生成的文本作为外部线索也是如此。线性探测(linear probing)揭示,虽然生成的文本几乎不携带量化的痕迹,但内部表示包含清晰的、方法特定的指纹(fingerprints)。通过训练,模型学会通过比较识别严重退化的输出(例如4-bit模型的输出),但仍无法从单个输出中做到这一点。跨量化级别联合训练的共享LoRA成功读取了内部指纹,但在未见过的量化方法上失败,仅将方法特定的指纹映射到标签。虽然外部自我观察可以在某些人类病感失认案例中恢复意识,但我们的结果表明,在LLM中实现这种意识的更有希望的途径可能在于其内部表示。我们的结果凸显了LLM自我监控泛化性的根本限制。

英文摘要

Can LLMs recognize degradation in their own computational substrate? Inspired by anosognosia, a neurological condition in which patients fail to recognize impairments in their own abilities, we investigate whether LLMs can recognize degradation in their computational substrate induced by quantization. We first show that existing models fail to self-report their quantization state, even when provided with their own generated text as an external cue. Linear probing reveals that, while generated text carries almost no trace of quantization, internal representations contain clear, method-specific fingerprints. Through training, models learn to identify severely degraded outputs such as those of 4-bit models by comparison, yet still fail to do so from a single output. A shared LoRA trained jointly across quantization levels succeeded in reading out internal fingerprints, but fails on unseen quantization methods, merely mapping method-specific fingerprints to labels. Whereas external self-observation can restore awareness in some cases of human anosognosia, our results suggest that the more promising route to enabling such awareness in LLMs may lie in their internal representations. Our results highlight fundamental limits of generalizability to LLM self-monitoring.

Comments9 main pages with appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑