arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI对语言的表征中,虚假与不可能是不同的方向

Falsehood and Impossibility Are Different Directions in an AI's Representation of Language

Yoon Pyo Lee

arXiv 2608.12852首次发表:更新:

发表机构

University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过对Gemma 3 4B IT模型的激活分析,发现其内部能区分语言的不可能性与虚假,但将偶然虚假混同于矛盾,且不可能性表征与真值、语义异常表征方向不同。

AI 中文摘要

语言可以描述虚假的事态和根本不可能存在的事态,AI模型是否在内部区分这些失败尚不明确。本研究对多模态开放权重模型Gemma 3 4B IT开展探索性激活研究,使用来自17个哲学类别的85个提示词,以及15个主题的主题匹配模态集,每个主题分别以真命题、偶然虚假命题、不可能主张、语义异常命题、必然虚假命题的形式呈现。在模型的回答中,它将偶然虚假与矛盾混为一谈,给15个虚假陈述中的12个贴上“矛盾”标签;但其激活呈现出不同模式:线性真值探针可区分不可能陈述与真实陈述(AUC=0.93),但无法区分不可能陈述与虚假陈述(AUC=0.20);在保留的主题类别上评估的不可能性探针,可区分必然虚假与偶然虚假,AUC达1.00,在第15层达到峰值,平衡准确率为0.97(经Bonferroni校正的P=0.018)。真值方向与不可能性方向接近正交,而不可能性方向与语义异常方向部分重叠但仍可区分,同一层的稀疏自编码器特征重复了这一几何结构:对不可能性有选择性的特征也会激活异常句子,但很少激活偶然虚假命题。在该模型的激活空间中,必然虚假并非偶然虚假的极端情况,而是更接近实验定义的语义异常类别;这种表征接近性并不意味着不可能陈述本质上无意义,来自这个小型模型的这些相关性观察,为一个古老的哲学区分提供了实证注脚。

英文摘要

Language can describe states of affairs that are false and states of affairs that could not be the case at all. Whether an AI model internally distinguishes these failures remains unclear. I report an exploratory activation study of the multimodal open-weight model Gemma 3 4B IT using 85 prompts from 17 philosophical families and a topic-matched modality set of 15 topics, each expressed as a truth, contingent falsehood, improbable claim, semantic anomaly, and necessary falsehood. In its answers, the model conflates contingent falsehood with contradiction, labeling 12 of 15 false statements "contradiction." Its activations show a different pattern. A linear truth probe separates impossible from true statements (AUC 0.93) but not impossible from false statements (AUC 0.20). An impossibility probe evaluated on held-out topic families separates necessary from contingent falsehood at AUC 1.00, peaking at layer 15 with balanced accuracy 0.97 (Bonferroni-adjusted P=0.018). The truth and impossibility directions are close to orthogonal, whereas the impossibility direction partially overlaps a semantic anomaly direction while remaining distinguishable from it. Sparse autoencoder features at the same layer repeat this geometry. Features selective for impossibility also fire on anomalous sentences but rarely on contingent falsehoods. In this model's activation space, necessary falsehoods are not extreme cases of contingent falsehood but lie closer to the experimentally defined category of semantic anomaly. This representational proximity does not imply that impossible statements are intrinsically meaningless. These correlational observations from one small model offer an empirical footnote to an old philosophical distinction.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑