arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18816cs.CLcs.AI

大语言模型会产生“电气海市蜃楼”式的幻觉吗?

Do Large Language Models Hallucinate Electric Fata Morganas?

Kristina Šekrst

AI总结:

本文探讨大语言模型幻觉的哲学意义,通过两项实证研究分析其成因,结合相关理论指出模型的情感自我陈述属于幻觉范畴,未来机器意识或因与高级幻觉无法区分而无法被认知。

AI中文摘要:

AI幻觉——即生成编造的、无法验证的或与源材料矛盾的输出——通常被视为需要处理的工程缺陷。本文认为,它们在机器意识问题上也具有哲学意义。我们研究了大语言模型幻觉的已知原因,如源-目标 divergence(差异)、训练与推理间的偏差、过拟合,并开展了两项实证研究。第一项研究中,我们将不同代次的GPT模型应用于不同温度设置下的模糊事实问题,发现更高的温度会产生看似合理但错误的答案,而更低的温度则会得到事实准确的答案;使模型表现出创造性或自发性、更可能通过智能行为测试的采样参数,同时也会提高其幻觉率。第二项研究中,我们考察了一个在百科全书式数据上训练的仅编码器模型,该模型回答同类问题时符合事实且无修饰,表明幻觉源于接触主观且社会多样化的训练数据,而非任何认知能力的发展。结合图灵、塞尔的中文屋、框架问题,以及维纳和阿什比的控制论传统,我们主张,模型对情感或感知的自我陈述属于幻觉的定义范畴,且任何未来机器意识的出现可能在认识论层面无法被感知,因为它将与足够先进的幻觉无法区分。

英文摘要:

AI hallucinations - that is, outputs which are made up, cannot be verified, or contradict the source material - are generally regarded as an engineering flaw to be dealt with. This paper contends that they also have philosophical significance when it comes to the question of machine consciousness. We examine the known causes of hallucinations in large language models - such as source-target divergence, discrepancies between training and inference, and overfitting - and we present two empirical investigations. In the first, we apply successive generations of the GPT model to ambiguous factual questions under different temperature settings, finding that higher temperatures result in plausible but incorrect answers while lower temperatures lead to factually accurate ones. The sampling parameters that cause a model to seem creative or spontaneous and thus more likely to pass behavioral tests of intelligence are the same ones that increase its hallucination rate. In the second, we look at an encoder-only model that has been trained on encyclopedic data and which answers questions of the same type factually and without embellishment, indicating that hallucinations are due to exposure to subjective and socially diverse training data rather than to the development of any cognitive ability. Using references to Turing, Searle's Chinese Room, the frame problem, and the cybernetic tradition of Wiener and Ashby, we claim that a model's self-reports of emotion or sentience come within the definition of hallucination, and that any future occurrence of machine consciousness might remain epistemically inaccessible since it would be indistinguishable from a sufficiently advanced hallucination.

补充信息

↑