arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

解耦不等于识别:下一词预测中的监督证据学习

Decoupling Is Not Identification: Supervised Evidential Learning in Next-Token Prediction

Ge Wang

arXiv 2609.26268首次发表:更新:

发表机构

Rensselaer Polytechnic Institute(伦斯勒理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出ENTOP框架,在字符级《白鲸》上通过显式监督证据学习改进下一词预测中的支持估计,优于隐式方法,并强调最小证据协议以验证集中度的识别、校准与效用。

AI 中文摘要

下一词概率表示模型预测了什么,而非其背后有多少训练支持。Dirichlet 头通过分离均值 $m$ 与集中度 $S$ 来表示这一区别,但解耦并不能识别 $S$ 的含义。在此,我们提出一个证据性下一词预测(ENTOP)框架,在字符级《白鲸》上审计这一差距,使用精确的 8-gram 计数作为可复现的词汇支持标签,并从 20% 的上下文类型中扣留计数回归。标准隐式证据训练在保留标签类型上除了置信度外几乎不携带计数信号(部分 Spearman $\rho = 0.001 \pm 0.014$),而显式监督则具有泛化能力($\rho = 0.201 \pm 0.010$;配对胜率 $= 0.822 \pm 0.021$)。CE 预测熵对未见 8-gram 的判别接近随机(AUROC $= 0.490 \pm 0.004$),而监督空虚度达到 $0.772 \pm 0.003$,与索引 CE 表示基线($0.769$)相当,但低于同义反复的语料库神谕($1.000$)。最长后缀或表示距离分层均无法解释摊销成功的所在。在 digamma 目标下增加计数权重仅通过牺牲预测来改善支持拟合。常数预测器在自然对数 RMSE 上获胜,且空虚度并未改善错误延迟。这些结果促使提出最小证据协议——置信度控制、配对、保留标签、常数基线和决策测试——并表明集中度可以通过识别测试,但在校准和效用上失败。

英文摘要

A next-token probability says what a model predicts, not how much training support lies behind it. A Dirichlet head can represent this distinction by separating mean $m$ from concentration $S$, but decoupling does not identify what $S$ means. Here we propose an Evidential Next-Token Prediction (ENTOP) framework to audit this gap on character-level Moby-Dick, using exact 8-gram count as a reproducible lexical-support label and withholding count regression from 20% of context types. Standard implicit evidential training carries essentially no count signal beyond confidence on held-out-label types (partial Spearman $ρ= 0.001 \pm 0.014$), whereas explicit supervision generalizes ($ρ= 0.201 \pm 0.010$; matched-pair win $= 0.822 \pm 0.021$). CE predictive entropy is at chance for unseen 8-grams (AUROC $= 0.490 \pm 0.004$), while supervised vacuity reaches $0.772 \pm 0.003$, comparable with an indexed CE-representation baseline ($0.769$) but below the tautological corpus oracle ($1.000$). Neither longest-suffix nor representation-distance strata explain where amortization succeeds. Increasing count weight under the digamma objective improves support fit only by sacrificing prediction. A constant predictor wins natural log-RMSE, and vacuity does not improve error deferral. These results motivate a minimum evidence protocol---confidence control, matched pairs, held-out labels, a constant baseline, and a decision test---and show that concentration can pass identification while failing calibration and utility.

Comments5 pages, 3 figures, and 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑