发表机构
Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算学部)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究证明给定分区控制语言模型中证据权重,通过去重聚合和误差界分析,实验显示分区变化导致依赖检查点的决策偏移。
AI 中文摘要
检索到的记录是呈现单元;给定的分区决定哪些记录作为一个证据贡献进入语言模型。我们刻画了不变的分组内容状态,该状态在保留互补的规范内容的同时移除组内副本,表明相等的组计数可以编码不同的证据状态,并推导出一个尖锐的内容感知分区误差界。给定一个分区,我们的生成前表示在组内去重并聚合内容,并限制每个组的贡献。在104,402次试验和6个公共检查点上,一项中心自然文本干预发现,内容固定的错误拆分增加了10.27-32.66个百分点,错误合并减少了9.13-31.79个百分点;一个匹配的六槽对照在所有16个单元格中保持正向方向。在一个新的48项受控活动面板中,更改给定分区在所有四个模型中产生可测量的、依赖检查点的决策偏移,平衡镜像设计暴露了显著的顺序交互。综合来看,理论和实验确立了给定分区作为可控的生成前表示变量,并刻画了其依赖检查点的行为效应。
英文摘要
Retrieved records are presentation units; a supplied partition determines which records enter a language model as one evidential contribution. We characterize the invariant group-content state that removes within-group copies while retaining complementary canonical content, show that equal group counts can encode different evidence states, and derive a sharp content-aware partition-error bound. Given a supplied partition, our pre-generation representation deduplicates and aggregates content within groups and bounds each group's contribution. Across 104,402 trials and 6 public checkpoints, a central natural-text intervention finds that content-fixed false splits add 10.27-32.66 percentage points and false merges remove 9.13-31.79 points; a matched six-slot control retains the positive direction in all 16 cells. In a new 48-item controlled campaign panel, changing the supplied partition produces measurable, checkpoint-dependent decision shifts across all four models, and the balanced mirror design exposes substantial order interactions. Together, the theory and experiments establish the supplied partition as a controllable pre-generation representation variable and characterize its checkpoint-dependent behavioral effects.
Comments28 pages, 5 figures