通过注意力头重新加权实现大语言模型的数据高效适应
Data-Efficient Adaptation of LLMs via Attention Head Reweighting
AI总结:
针对大语言模型在有限数据学习中的难题,提出注意力头重新加权(AHR)方法,通过为每个注意力头学习单个标量来适应新任务,大幅减少需学习的参数,实验表明该方法在有限样本学习时优于标准基线,且权重易解释。
AI中文摘要:
在安全等标记示例稀缺的领域,从有限数据中有效学习至关重要。大语言模型(LLMs)已展现出数据高效学习的能力,尤其是通过参数高效适应方法,但面对困难任务的少量样本时仍有挑战。为此提出注意力头重新加权(AHR)方法,通过为每个注意力头学习单个标量来使LLMs适应新的文本分类任务,利用注意力头功能专业化大幅减少需学习的参数数量。在多种开源文本分类数据集上的实验表明,AHR在从有限样本学习时能超越LoRA等标准基线,虽可训练参数少200 - 1000倍,但仅修改约0.0001%的模型参数。此外,学习到的权重易于解释,有助于更好理解LLMs上下文学习能力的机制和注意力头。
英文摘要:
Learning effectively from limited data is critical in domains like security where labeled examples are scarce. Large language models (LLMs) have demonstrated some capabilities for data-efficient learning, especially through parameter-efficient adaptation methods, but continue to struggle when faced with few samples for difficult tasks. To meet this challenge, we propose Attention Head Reweighting (AHR), a data-efficient method that adapts LLMs to new text-classification tasks by learning only a single scalar per attention head. This drastically reduces the number of parameters that need to be learned by making use of the functional specialization of individual attention heads. Experiments on diverse open-source text classification datasets show that AHR can outperform standard baselines like LoRA when learning from limited samples, despite having 200-1000x fewer trainable parameters, as our AHR only modifies ~0.0001% of the model's parameters. In addition, our learned weights are easy to interpret and can be analyzed to better understand the mechanisms and attention heads responsible for in-context learning abilities in LLMs.