右侧重置:通过前缀移除进行分块
Right Reset: Chunking by Prefix Removal
AI总结:
本研究提出右侧重置(RR)方法,通过前缀移除探测将因果语言模型的上下文依赖转化为边界信号,在文本分块任务上优于基线方法,减少局部输出干扰。
AI中文摘要:
从因果语言模型中移除左侧上下文会揭示出一种有用的边界:模型处理相同右侧标记时几乎没有变化的边缘。我们将这一观察转化为前缀移除探测,并引入右侧重置(Right Reset, RR),该方法用于衡量右侧隐藏状态轨迹的保留情况。动态规划将 RR 边缘分数转换为可变长度的分块。在通过删除主题相似记录的分隔符和布局后拼接形成的扁平化文本上,RR 恢复了 47.7% 的原始记录作为干净单元,而最强的常规基线 BGE 嵌入边界(未进行特定任务模型训练)仅恢复了 25.9%。在渲染和光学字符识别(OCR)后,该增益仍然存在。来自同一 Qwen3-4B 层的被动分数以及相同规模指令模型的直接提示在扁平化记录上表现明显更差。在六个语言模型中,与未选中的候选边缘相比,RR 选择的分割处始终经历更少的局部输出干扰。观察到的标记似然比读数在某些架构中具有竞争力,这表明核心贡献在于该干预措施:当表面结构较弱时,上下文依赖本身可以提供边界信号。
英文摘要:
Removing the left context from a causal language model reveals a useful kind of boundary: an edge where the model processes the same right-hand tokens with little change. We turn this observation into prefix-removal probing and introduce Right Reset (RR), which measures preservation of the right-hand hidden-state trajectory. A dynamic program converts RR edge scores into variable-length chunks. On flattened text formed by concatenating topically similar records after deleting their separators and layout, RR recovers 47.7% of the original records as clean units, versus 25.9% for a BGE embedding boundary, the strongest tested conventional baseline without task-specific model training. The gain persists after rendering and OCR. Passive scores from the same Qwen3-4B layer and direct prompting of a same-scale instruction model perform substantially worse on flattened records. Across six language models, RR-selected cuts also undergo consistently less local output disruption than unselected candidate edges. An observed-token likelihood-ratio readout is competitive in some architectures, indicating that the central contribution is the intervention: context dependence itself can provide a boundary signal when surface structure is weak.