arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20421cs.CYcs.AI

大型语言模型的六大误解:一个极简模型与诊断分类法

Six misconceptions about large language models: A minimal model and diagnostic taxonomy

  • Yonsei University(延世大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhicheng Lin

AI总结:

该研究提出以四组区分为核心的LLM极简模型,诊断六大误解,应用于出版商AI政策案例,为纠正民间理论错误提供诊断工具。

AI中文摘要:

大型语言模型(LLMs)现已嵌入科学、教育和治理工作流程,相关争论围绕其能力、机制与影响展开,但这些争论仍受持续存在的民间理论——即指导态度与行动的直觉性、非正式解释模型——的结构化影响。诸如“只是自动补全”“随机鹦鹉”“互联网平均值”等紧缩性口号,以及“涌现智能体”“原心智”等拟人化框架,各自捕捉了当前系统的真实特征,但将这些特征误判为整体。本文提出了一个以四组区分关系为核心的LLM系统极简工作模型:预训练与部署系统的区分、学习分布与特定样本的区分、参数记忆、上下文记忆与外部记忆的区分,以及任务能力与智能体的区分。该模型用于诊断关于LLMs的六大误解:下一个词元预测、均值回归、训练数据复述、模型记忆、对齐与理解。对每个误解,分析会明确其正确之处、混淆的区分关系,以及对能力评估、系统设计和治理的启示。将该框架应用于出版商AI政策作为治理案例研究,结果显示政策语言如何混淆这些区分,以及此类错误如何被纠正。该模型通过将LLMs视为话语与任务表现的模拟器,而非鹦鹉-心智二元体,提供了一套诊断工具,用于定位并纠正这些民间理论持续存在的错误。

英文摘要:

Large language models (LLMs) are now embedded in scientific, educational, and governance workflows, with debates centering on their capabilities, mechanisms, and impacts. Yet these debates remain structured by persistent folk theories--intuitive, informal explanatory models that guide attitudes and actions. Deflationary slogans ("just autocomplete," "stochastic parrots," and "average of the internet") and anthropomorphic framings ("emergent agents" and "proto-minds") each capture genuine features of current systems but mistake those features for the whole. This Perspective proposes a minimal working model of LLM-based systems centered on four distinctions: between pretraining and deployed systems; between the learned distribution and particular samples; among parametric, contextual, and external memory; and between task competence and agency. The model is used to diagnose six misconceptions about LLMs: next-token prediction, regression to the mean, training-data regurgitation, model memory, alignment, and understanding. For each, the analysis identifies what the misconception gets right, which distinctions it conflates, and what follows for capability evaluation, system design, and governance. Applied to publisher AI policies as governance case studies, the framework shows both how policy language can conflate these distinctions and how such errors can be corrected. The model thereby avoids the parrot-mind binary by treating LLMs as simulators of discourse and task performance, offering a diagnostic toolkit for locating and correcting the errors these folk theories perpetuate.

补充信息

↑