arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12756cs.CLcs.LG

ReconSpan:重建引导的自适应隐式分词

ReconSpan: Reconstruction-Guided Adaptive Latent Tokenization

  • Cornell University(康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

Lixing Li

AI总结:

本文提出ReconSpan,一种重建引导的自适应隐式分词方法,可将文本划分为特定块并保留前缀码作为隐式token,在匹配平均长度下其边界比随机边界保留更多文本,能可靠传递主题信息但难以提取精确细节。

AI中文摘要:

自适应隐式分词将细粒度输入映射为更短的连续表示序列,该序列与依赖输入的文本块相关联。本文提出ReconSpan,它将文本划分为若干块,每个块都可由反向解码器从单个上下文前缀码重建,且每个块保留一个此类码作为隐式token。在形成块时应用重建准则,使得训练后的自动编码器能产生平均块长度为6.5至12.2的结果。在匹配的平均长度下,重建引导的边界比随机边界保留更多文本;从所得隐式序列中,读者可可靠恢复主题信息,但难以提取精确细节。

英文摘要:

Adaptive latent tokenization maps a fine-grained input to a shorter sequence of continuous representations associated with input-dependent spans. We introduce ReconSpan, which divides text into chunks that a backward decoder can reconstruct from a single contextual prefix code and retains one such code as the latent token for each chunk. The reconstruction criterion is applied when chunks are formed, allowing one trained autoencoder to produce average chunk lengths from 6.5 to 12.2. At matched average length, reconstruction-guided boundaries preserve more text than random boundaries. Readers of the resulting latent sequence recover topic information reliably but struggle to extract exact details.

补充信息

↑