arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

人类可读文本对于大语言模型的有效微调是否必要?

Is Human-Readable Text Necessary for Effective LLM Fine-Tuning?

Jinhao Zhang, Zeyu Liu, Zicheng Yan, Yunquan Zhang, Daning Cheng, Song Tang

arXiv 2609.35868首次发表:更新:

发表机构

Beijing University of Posts and Telecommunications; Institute of Computing Technology, Chinese Academy of Sciences; University of Science and Technology of China; University of Shanghai for Science and Technology(北京邮电大学; 中国科学院计算技术研究所; 中国科学技术大学; 上海理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出 DASA 方法,利用激活梯度反馈优化连续嵌入,无需人类可读文本即可实现与大语言模型微调相当或更优的性能,并在多个基准上超越现有方法,同时显著加速训练。

AI 中文摘要

人类可读性对于大语言模型的有效微调是否必要?我们研究了模型条件下的训练表示能否在不需要人类可读文本形式的情况下保持或提升适应效用。我们提出了期望更新对齐合成数据(DASA),该方法利用冻结参考模型的激活梯度反馈来指导连续合成输入嵌入的优化。受激活梯度在局部风险降低中作用的启发,DASA 针对有用的适应更新而非源文本重建或语言流畅性。生成的嵌入直接用于下游微调;离散令牌投影仅用于定性检查。在 Llama 和 Qwen 家族的六个模型上进行的实验,参数范围从 1B 到 32B,覆盖了涵盖知识、数学推理、代码生成和常识推理的六个基准。在匹配的 LoRA 适应设置下,DASA 达到了与源自然语言数据相当的性能,并在多种配置中超越之,同时在大多数比较中优于 GRADMM。进一步的实验涵盖了通用领域和任务专用源数据。在所评估的合成设置下,DASA 相对于 GRADMM 提供了 3.6 至 4.9 倍的加速,且峰值 GPU 内存相当。

英文摘要

Is human readability necessary for effective fine-tuning of large language models? We investigate whether model-conditioned training representations can preserve or improve adaptation utility without requiring a human-readable textual form. We propose Desired-Update-Aligned Synthetic Data (DASA), which uses activation-gradient feedback from a frozen reference model to guide the optimization of continuous synthetic input embeddings. Inspired by the role of activation gradients in local risk reduction, DASA targets useful adaptation updates rather than source-text reconstruction or linguistic fluency. The resulting embeddings are used directly for downstream fine-tuning; discrete token projections are employed only for qualitative inspection. Experiments on six models from the Llama and Qwen families, ranging from 1B to 32B parameters, cover six benchmarks spanning knowledge, mathematical reasoning, code generation, and commonsense reasoning. Under matched LoRA adaptation settings, DASA achieves performance comparable to the source natural-language data and surpasses it in multiple configurations, while outperforming GRADMM in most comparisons. Further experiments cover general-domain and task-specialized source data. Under the evaluated synthesis settings, DASA provides a $3.6$--$4.9\times$ speedup over GRADMM with comparable peak GPU memory.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑