arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型中的上下文绑定容量

In-Context Binding Capacity in Language Models

Manas Venkata Sai Ravulapalli, Samrath Singh Chadha

arXiv 2609.30634首次发表:更新:

发表机构

Efficient Computation Inc.(高效计算公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过连续回忆曲线和阈值扫描测量语言模型的上下文绑定容量,发现容量随规模呈幂律增长,并分析了干扰、训练及测量标准对容量评估的影响,为工作记忆和指令遵循研究提供基线。

AI 中文摘要

语言模型在忘记哪个值属于哪个实体之前,能回忆起多少个赋值?我们使用12个参数规模在30亿及以下的模型的连续回忆曲线,以及对30个参数规模高达120亿的开放模型的阈值扫描来测量这一限制。在连续曲线上,回忆降至随机水平一半时的负载遵循 $K_{50}=cN^{\alpha}$,其中 $\alpha=0.820$,$R^2=0.73$。更广泛的扫描显示,与预训练配方相关的八倍范围,尽管在控制规模后,连续曲线未显示可检测的配方效应,且拟合中很少有现代模型。我们推导了为什么干扰可以通过降低单绑定回忆来降低测量容量,即使负载依赖的回忆曲线不变。直接任务训练也超过了外推的零样本定律,但不同的测量标准阻止将这种比较解释为容量提升。其形成时间在两个独立的代码库中遵循幂律形式,取决于成功的运行。总之,这些结果表征了模型查询接口处的容量。联合回忆的界限和政策错误的分解将这一测量与工作记忆和指令遵循联系起来,而不将回忆视为对齐的度量。受控任务还为测试绑定限制是否约束世界状态跟踪提供了基线;本实验未测量状态更新或下游迁移。

英文摘要

How many assignments can a language model recall before it loses track of which value belongs to which entity? We measure this limit using continuous recall curves for 12 models at or below 3B parameters and a threshold sweep over 30 open models up to 12B. On the continuous curves, the load at which recall falls halfway to chance follows $K_{50}=cN^α$, with $α=0.820$ and $R^2=0.73$. The broader sweep shows an eightfold range associated with pretraining recipe, although the continuous curves show no detectable recipe effect after controlling for scale, with few modern models in the fit. We derive why interference can lower measured capacity by reducing single-binding recall even when the load-dependent recall profile is unchanged. Direct task training also exceeds the extrapolated zero-shot law, but different measurement criteria prevent interpreting that comparison as a capacity gain. Its formation times follow a power-law form in two independent codebases, conditional on runs that succeed. Together, these results characterize capacity at the model's query interface. Bounds on joint recall and a decomposition of policy errors connect this measurement to working memory and instruction following, without treating recall as a measure of alignment. The controlled task also provides a baseline for testing whether binding limits constrain world-state tracking; the present experiments do not measure state updates or downstream transfer.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑