arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言中的联合嵌入预测架构悖论:语言替代方案的几何结构

The JEPA Paradox in Language: The Geometry of Linguistic Alternatives

Anh Trac Duc Dinh, Khang Nhat Hoang Vo

arXiv 2607.23531首次发表:更新:

发表机构

Center for AI Research (CAIR), VinUniversity; Ho Chi Minh City University of Technology (HCMUT); Mohamed bin Zayed University of Artificial Intelligence(人工智能研究中心(CAIR),文大大学; 胡志明市理工大学(HCMUT); 穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究语言中JEPA未成为文本编码器标准目标的问题,通过可预测性等三个条件形式化平方误差潜在预测与语言条件结构的不匹配,实验揭示相关序列及模式,表明文本兼容的JEPA目标应保留多个合理完成而非压缩为单个潜在点。

AI 中文摘要

联合嵌入预测架构(JEPA)在图像、视频和音频方面很有效,但确定性的JEPA式潜在预测尚未成为文本编码器的标准目标。我们认为这种差距反映了平方误差潜在预测与语言条件结构之间的不匹配。关键要求是条件集中:给定上下文和目标位置,目标表示应位于单个有意义的点附近。局部图像预测通常通过空间连续性满足这一点,而掩码文本可以允许多个有效的令牌或跨度完成,其表示不需要共享连贯的中心。我们通过可预测性、不坍缩和低条件方差这三个条件来形式化这种不匹配,并展示它们的失败如何在文本中产生质心退化和坍缩压力。匹配的I-JEPA和T-JEPA实验揭示了预测序列:互信息饱和和目标方差升高先于训练-验证不稳定性、有效秩退化、余弦坍缩和下游转移不佳。相同模式出现在五个独立的数据种子中,表明这不是采样伪像。这些结果不排除对语言的预测学习;它们表明与文本兼容的JEPA目标必须保留多个合理的完成,而不是将它们压缩到单个潜在点。

英文摘要

Joint-Embedding Predictive Architectures (JEPAs) are effective for images, video, and audio, yet deterministic JEPA-style latent prediction has not become a standard objective for text encoders. We argue that this gap reflects a mismatch between squared-error latent prediction and the conditional structure of language. The key requirement is conditional concentration: given a context and target location, the target representation should lie near a single meaningful point. Local image prediction often satisfies this through spatial continuity, whereas masked text can admit multiple valid token or span completions whose representations need not share a coherent center. We formalize this mismatch through three conditions---predictability, non-collapse, and low conditional variance---and show how their failure creates centroid degeneracy and collapse pressure in text. Matched I-JEPA and T-JEPA experiments reveal the predicted sequence: mutual-information saturation and elevated target variance precede train--validation instability, effective-rank degeneration, cosine collapse, and poor downstream transfer. The same pattern appears across five independent data seeds, indicating that it is not a sampling artifact. These results do not rule out predictive learning for language; they show that text-compatible JEPA objectives must preserve multiple plausible completions rather than compress them into a single latent point.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑