层感知微调:通过大语言模型的三阶段功能分割
Layer-Informed Fine-Tuning via Three-Stage Functional Segmentation of LLMs
浏览论文内容
中文总结 AI 辅助
本研究提出层感知微调(LIFT)方法,通过敏感性分析识别大语言模型在概念化、推理和文本化三阶段中的关键功能层,仅更新这些层以实现高效微调,实验证明该方法能加速训练并显著提升模型性能。
中文摘要 AI 辅助
近年来,大型语言模型(LLM)在推理任务上的表现十分出色,甚至在各种基准测试中超越了人类能力。然而,学术界对于LLM的结构和内部参数如何逐步解决复杂推理问题仍缺乏清晰的理解。在本研究中,我们考察了LLM在跨语言材料上的推理过程,并提出假设:LLM的各层在概念化、推理和文本化方面表现出结构化的分工。基于这一假设,我们引入了一种使用敏感性分析的瓶颈识别机制,以定位特定任务中最关键的功能阶段。利用这一见解,我们提出了一种新方法,即层感知微调(LIFT),通过仅选择性地更新这些功能关键层来实现高效且有效的微调。随后,我们进行了大量实验,表明LIFT方法不仅加速了训练过程,而且显著提升了模型性能。
英文摘要
In recent years, the performance of large language models (LLMs) on reasoning tasks has been remarkable, even surpassing human capabilities on various benchmarks. However, there remains a lack of clear understanding in the academic community regarding how the structure and internal parameters of LLMs progressively solve complex reasoning problems. In this study, we investigate the inference process of LLMs on cross-linguistic materials and propose the hypothesis that LLM layers exhibit a structured division of labor across conceptualization, reasoning, and textualization. Based on this hypothesis, we introduce a bottleneck identification mechanism using sensitivity analysis to pinpoint the most critical functional stage for a specific task. Leveraging this insight, we propose a novel approach, Layer-Informed Fine-Tuning (LIFT), which achieves efficient and effective fine-tuning by selectively updating only these functionally critical layers. We then conduct extensive experiments to show that the LIFT method not only accelerates the training process but also significantly improves model performance.
发表机构
- Tsinghua University(清华大学)
- Microsoft Research Asia(微软亚洲研究院)
- Shanghai Qi Zhi Institute(上海期智研究院)
机构由 AI 辅助整理,请以论文原文为准。