AI 中文总结
PARALLEL是一种前额叶对齐的强化启发式语言模型学习方法,通过分配样本相关更新强度提升适配效率,在多任务上保留高性能且适配轨迹更稳定,支持高效稳定的部署后流式适配。
AI 中文摘要
近期的语言模型在各类任务上表现出色,但传统适配方法会对所有训练样本进行均匀更新,忽略了各样本的局部更新收益。我们提出PARALLEL,一种前额叶对齐的强化启发式语言模型学习方法。受目标相关控制与不确定性相关控制的互补作用启发,PARALLEL将这两类信息表示为独立的控制器信号,并与当前模型表示相结合。一个受强化启发的控制器利用即时效用-成本反馈分配样本相关的更新强度,因此PARALLEL能学习何时以及以何种强度适配每个样本,优先选择有益更新同时限制不必要的参数变化。PARALLEL比选择性基线更高效地利用可用更新,同时保留了全适配(Full-adaptation)性能的94.1%至99.2%。除了多选推理任务,在XSum和CNN/DailyMail上的实验显示,PARALLEL保留了全适配方法达到的ROUGE-1和ROUGE-2分数的96.9%至98.6%,以及对应ROUGE-L分数的98.8%至98.9%。当在相同的累计适配时间或GPU能耗下进行比较时,PARALLEL在代表性运行中实现了比全适配方法更高的ARC准确率,并展现出更稳定的后期适配轨迹。这些结果表明,学习何时以及以何种强度更新每个样本,可支持稳定高效的部署后流式适配,同时避免不必要的更新。
英文摘要
Recent language models achieve strong performance across a variety of tasks, but conventional adaptation applies updates uniformly across training samples regardless of their local update benefit. We propose PARALLEL, a prefrontal-aligned reinforcement inspired approach for language-model learning. Inspired by the complementary roles of goal-related and uncertainty-related control, PARALLEL represents these forms of information as separate controller signals and combines them with the current model representation. A reinforcement-inspired controller assigns sample-dependent update intensity using immediate utility-cost feedback. PARALLEL therefore learns when and how strongly to adapt to each sample, prioritizing beneficial updates while limiting unnecessary parameter changes. PARALLEL uses available updates more efficiently than selective baselines while retaining 94.1--99.2\% of Full-adaptation performance. Beyond multiple-choice reasoning, experiments on XSum and CNN/DailyMail show that PARALLEL retains 96.9--98.6\% of the ROUGE-1 and ROUGE-2 scores achieved by Full adaptation and 98.8--98.9\% of the corresponding ROUGE-L scores. When compared at the same cumulative adaptation time or GPU energy, PARALLEL achieves higher ARC accuracy and exhibits a more stable late-stage adaptation trajectory than Full adaptation in the representative run. These results show that learning when and how strongly to update each sample supports stable and efficient post-deployment stream adaptation while avoiding unnecessary updates.
Comments8 pages, 3 figures, and 5 tables