机器的内部时钟:大型语言模型(LLMs)是否会共享人类的时间错觉?
The Machine's Internal Clock: Do LLMs Share Human Temporal Illusions?
AI总结:
本研究构建含5种错觉的6684个叙事对基准,发现人类仅在2种错觉中偏好预期场景,而14个LLMs在4种错觉中与文献预测一致,其一致性源于心理学研究检索而非类人时间偏差。
AI中文摘要:
人类对时间的感知具有主观性,已有充分记录的时间错觉表明,大脑依赖语境和关系线索来判断时长,而非直接追踪流逝的时间。先前研究已通过视觉和听觉刺激证实了这些效应。现有的大型语言模型(LLMs)时间感知评估聚焦于事件时长估计或多步时间推理。本研究使用包含5种错觉的6684个叙事对组成的新基准,探究仅书面叙事是否能唤起人类的时间错觉。研究发现,60名人类读者仅在5种错觉中的2种中偏好预期场景,这2种错觉的操纵在文本中直接可见,无需读者内部模拟时长。在同一基准上评估14个LLMs后,惊讶地发现模型在5种错觉中的4种里选择了文献预测的场景,与人类行为存在差异。推理轨迹显示,约70%的响应明确唤起心理学研究,表明这种一致性与检索已发表的发现而非类人时间偏差相符。
英文摘要:
Human perception of time is subjective. Well-documented temporal illusions show that the brain relies on context and relational cues for judging duration instead of tracking elapsed time directly. Prior studies established these effects with visual and auditory stimuli. Existing LLM evaluations of temporal perception focus on estimating event durations or multi-step temporal reasoning. In this work, we investigate whether written narratives alone can evoke human temporal illusions, using a new benchmark of 6,684 narrative pairs spanning five illusions. We find that human readers (60 participants) prefer expected scenarios in only two of the five illusions, those where the manipulation is directly visible in text rather than requiring readers to internally simulate duration. We evaluate 14 LLMs on the same benchmark. Surprisingly, we find that models pick the literature-predicted scenario across four of the five illusions, diverging from human behavior. Reasoning traces show that ~70% of responses explicitly evoke psychology research, suggesting that this alignment is consistent with retrieval of published findings rather than human-like temporal biases.