随机自回归学习
Stochastic Autoregressive Learning
- MIT(麻省理工学院)
- The Hebrew University(希伯来大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
该研究提出二元随机自回归PAC学习模型,分析基础、思维链、端到端三种监督下的学习样本量界,揭示其与确定性理论的本质差异,并通过d维逻辑函数验证结论的紧性。
中文摘要 AI 辅助
受大语言模型(LLM)通过从下一词元分布中迭代采样生成输出这一特性的启发,我们提出了一种用于二元随机自回归学习的PAC(概率近似正确)学习模型。该模型推广了Joshi等人在2025年计算学习理论会议(COLT 2025)上提出的确定性自回归学习框架。在我们的模型中,一个固定生成器为每个提示字符串分配一个伯努利下一词元分布。从输入提示开始,采样一个词元并追加到提示后;随后将同一生成器应用于扩展后的提示;该过程重复进行$M$步。我们考虑三种监督形式:基础单步样本、揭示长度为$M$的完整随机轨迹的思维链(CoT)样本,以及仅揭示长度为$M$的轨迹的最终词元的端到端(e2e)样本。针对一个生成器类别,我们研究在平方损失误差$\boldsymbol{\u03b5}$下,学习基础模型中的单步概率、以及CoT和e2e模型中的最终词元概率分别所需的最小样本量$m_{base}(\u03b5)$、$m_{CoT}(\u03b5)$、$m_{e2e}(\u03b5)$。\n我们证明,随机自回归学习与确定性理论存在本质差异。在尺度$\boldsymbol{\u03b5}$下,三种学习任务之间不存在普适的比较关系:$m_{CoT}/m_{base}$和$m_{e2e}/m_{CoT}$可以同时任意大于$M/\u03b5$——这是现有确定性结果的自然对应量。尽管如此,在调整尺度后,对于每个类别,尺度$\u03b5$下的CoT学习的样本量上界由尺度$\u03b5/M^2$下的基础学习样本量给出,而尺度$\u03b5$下的e2e学习的样本量上界(至多相差对数因子)由$(M/\u03b5) m_{CoT}(\u0398(\u03b5))$给出。这些依赖关系和尺度本质上是紧的。我们还通过研究模型中的$d$维逻辑函数补充了这些界的相关结论。
英文摘要
Motivated by LLMs, which generate outputs by iteratively sampling from next-token distributions, we introduce a PAC-learning model for binary stochastic autoregressive learning. This generalizes the deterministic autoregressive learning framework of Joshi et al., COLT 2025. In our model, one fixed generator assigns a Bernoulli next-token distribution to every prompt string. Starting from an input prompt, a token is sampled and appended to the prompt; the same generator is then applied again to this expanded prompt; this procedure is repeated for $M$ steps. Three forms of supervision are considered: base one-step samples, chain-of-thought (CoT) samples that reveal full random trajectories of length $M$, and end-to-end (e2e) samples that reveal only the final token of length $M$ trajectories. For a generator class, we study the minimum number of samples $m_{base}(\varepsilon),m_{CoT}(\varepsilon), m_{e2e}(\varepsilon)$, resp., required to learn the one-step probabilities in the base model, and the final-token probability in the CoT and e2e models, under squared loss error~$\varepsilon$. We show that stochastic autoregressive learning fundamentally differs from the deterministic theory. At scale $\varepsilon$, there is no universal comparison between the three learning tasks: both $m_{CoT}/m_{base}$ and $m_{e2e}/m_{CoT}$ can be made simultaneously arbitrarily larger than $M/\varepsilon$, the natural analogue for the existing deterministic results. Nevertheless, after altering scales, for every class, CoT learning at scale $\varepsilon$ is upper-bounded by base learning at scale $\varepsilon/M^2$, whereas e2e learning at scale $\varepsilon$ is upper-bounded, up to logarithmic factors, by $(M/\varepsilon) m_{CoT}(Θ(\varepsilon))$. These dependencies and scales are essentially tight. We complement these bounds by studying dimension $d$ logistic functions in our model.