发表机构
College of AI, Tsinghua University; Department of Electronic Engineering, Tsinghua University(清华大学人工智能学院; 清华大学电子工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对低资源LLM训练中数据模糊或不完整的问题,提出Prior-Guided Tuning(PGT)方法,引入Contrastive Prior Steering(CPS),在多个数据集上取得显著性能提升,验证了任务级自然语言先验作为辅助学习信号的有效性。
AI 中文摘要
大语言模型(LLM)在低资源训练数据模糊或不完整时往往表现不佳。任务级自然语言先验可在此类场景提供有用指导,但现有方法通常将这些先验视为输入上下文,而非训练过程中的学习信号。我们提出Prior-Guided Tuning(PGT,先验引导调优),这一训练视角将自然语言先验作为辅助学习信号用于低资源LLM训练。在此视角下,我们引入Contrastive Prior Steering(CPS,对比先验引导),其在保留原始监督目标不变的同时,添加基于正、负先验的辅助损失,以鼓励任务一致的学习并抑制看似合理但具误导性的替代方案。在AmbiMath、Jigsaw及MNLI/HANS上的实验显示,CPS相比普通微调与提示微调均有持续提升:在AmbiMath上,CPS达到97.6%的平均精确匹配准确率;在Jigsaw上,CPS相比标准微调将平均Macro F1提升9.5个百分点,且仅用1/10的实验训练数据就略优于全数据普通微调;在HANS上,CPS使LLaMA 3.1 8B和Qwen 2.5 7B的非蕴含准确率分别提升8.3和5.2个百分点,同时保持MNLI域内准确率相当。这些结果支持我们的核心主张:任务级自然语言先验可作为辅助学习信号为低资源LLM训练提供有用指导,我们的代码与数据将公开提供。
英文摘要
Large language models (LLMs) often struggle when low-resource training data are ambiguous or incomplete. Task-level natural-language priors can provide useful guidance in such settings, but existing approaches usually treat these priors as input context rather than as learning signals during training. We propose Prior-Guided Tuning (PGT), a training perspective that incorporates natural-language priors as auxiliary learning signals for low-resource LLM training. Under this perspective, we introduce Contrastive Prior Steering (CPS), which keeps the original supervised objective intact while adding positive and negative prior-conditioned auxiliary losses to encourage task-consistent learning and discourage plausible but misleading alternatives. Experiments on AmbiMath, Jigsaw, and MNLI/HANS show that CPS consistently improves over plain and prompt fine-tuning. On AmbiMath, CPS achieves 97.6% average exact-match accuracy. On Jigsaw, CPS improves average Macro F1 by 9.5 percentage points over standard fine-tuning, and with 1/10 of the experimental training data slightly exceeds full-data plain fine-tuning. On HANS, CPS improves non-entailment accuracy by 8.3 and 5.2 percentage points for LLaMA 3.1 8B and Qwen 2.5 7B, respectively, while maintaining comparable in-domain MNLI accuracy. These results support our central claim: task-level natural-language priors can provide useful guidance as auxiliary learning signals for low-resource LLM training. Our code and data will be publicly available.