TELLME:面向语言模型增强的测试增强学习
TELLME: Test-Enhanced Learning for Language Model Enrichment
浏览论文内容
中文总结 AI 辅助
本研究提出TELLME方法,将测试增强学习(TEL)原理与持续预训练(CPT)结合,缓解大语言模型领域自适应中数据与成本问题,在金融领域表现优于现有方法,长期记忆保留提升显著。
中文摘要 AI 辅助
持续预训练(CPT)已被广泛用作大语言模型领域自适应的方法,但它始终伴随挑战,如难以获取大规模领域特定数据集、计算成本高。本研究提出一种名为面向语言模型增强的测试增强学习(TELLME)的新方法以缓解这些问题。TELLME利用测试增强学习(TEL)原理,即在训练期间通过测验提升模型训练效率,将该原理与CPT结合,从而促进高效获取领域特定知识并增强长期记忆保留。实验结果表明,TELLME在金融领域的表现优于现有方法达23.6%,且长期记忆保留提升了9.8%。
英文摘要
Continual pre-training (CPT) has been widely adopted as a method for domain adaptation in large language models. However, CPT has consistently been accompanied by challenges, such as the difficulty of acquiring large-scale domain-specific datasets and high computational costs. In this study, we propose a novel method called Test-Enhanced Learning for Language Model Enrichment (TELLME) to alleviate these issues. TELLME leverages the TestEnhanced Learning (TEL) principle, whereby the model's training efficiency is improved using quizzes during training. It integrates this principle with CPT, thereby promoting efficient domain-specific knowledge acquisition and long-term memory retention. Experimental results demonstrate that TELLME outperforms existing methods by up to 23.6% in the financial domain and achieves a 9.8% improvement in long-term memory retention.
发表机构
- Korea Advanced Institute of Science and Technology(韩国科学技术院)
- Seoul National University of Science and Technology(首尔科学技术大学)
- National Security Research Institute(国家安全研究院)
机构由 AI 辅助整理,请以论文原文为准。