AI 中文总结
针对LLM-JEPA强对齐但预测不稳定的问题,提出ER-JEPA,通过情节回放存储并检索训练对提供额外监督,在多个数据集上稳定优于LLM-JEPA。
AI 中文摘要
大型语言模型(LLMs)擅长词元级生成,但可能学习到不良的抽象语义且缺乏全面的感知能力。LLM-JEPA通过联合嵌入预测架构(JEPA)对齐同一底层知识的不同视图来缓解这一问题。然而,强对齐并不一定能带来准确、稳定的预测。为解决此问题,我们提出ER-JEPA,它在LLM-JEPA中增加了一条情节回放路径。ER-JEPA将训练对存储在记忆中。在每一步中,它存储并检索相关数据以提供额外监督。这使得模型能够同时从当前批次和存储的训练对中学习,为词元预测和表示对齐提供额外监督。在多个数据集(NL-RX、GSM8K、Spider和NQ-Open)上的实验表明,ER-JEPA始终优于LLM-JEPA。
英文摘要
Large language models (LLMs) excel at token-level generation but may learn undesirable abstract semantics and lack comprehensive perception. LLM-JEPA mitigates this by aligning different views of the same underlying knowledge via a joint-embedding predictive architecture (JEPA). However, strong alignment does not necessarily lead to accurate, stable predictions. To address this, we propose ER-JEPA, which adds an episodic replay path to LLM-JEPA. ER-JEPA stores training pairs in a memory. At each step, it stores and retrieves relevant data to provide additional supervision. This enables learning from both the current batch and stored training pairs, providing additional supervision for token prediction and representation alignment. Experiments across multiple datasets (NL-RX, GSM8K, Spider, and NQ-Open) demonstrate that ER-JEPA consistently outperforms LLM-JEPA.
Comments20 pages, 15 figures, 6 tables