arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Token效用是选择条件化的:提示上下文与响应监督的耦合选择用于高效指令微调

Token Utility Is Selection-Conditioned: Coupled Selection of Prompt Context and Response Supervision for Efficient Instruction Tuning

Can Wu, Xinrui Chen, Ou Wu, Yi Du

arXiv 2609.22943首次发表:更新:

发表机构

Hangzhou Institute for Advanced Study; University of Chinese Academy of Sciences(杭州高等研究院; 中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对指令微调中提示上下文与响应监督选择状态不匹配问题,提出BRIDGE方法,通过共享交互代理和交替选择实现耦合选择,在数学推理、代码生成等任务上优于独立选择方法。

AI 中文摘要

高效的大语言模型(LLM)指令微调需要选择带有支持性提示上下文的响应监督。现有方法通常分别评估这两方面,导致评估与保留训练子集之间存在选择状态不匹配。BRIDGE(通过定向梯度引导的高效Token选择实现预算约束的响应-提示交互)通过一个共享的验证导向交互代理来捕捉选择条件化的Token效用,该代理在另一方的保留状态下评估每一方。预算约束的交替选择通过聚合当前对侧子集上的预计算交互来更新条件分数,从而协调保留子集。结构感知投影将条件响应分数转换为连贯的监督跨度。在三个模型家族中,BRIDGE在数学推理、代码生成和指令跟随方面总体上优于对比的选择方法。在数学推理中,其相对于独立选择的优势随压缩程度的增加而增长。

英文摘要

Efficient large language model (LLM) instruction tuning requires selecting response supervision with supporting prompt context. Existing methods typically value both sides separately, risking selection-state mismatch between valuation and retained training subsets. BRIDGE (Budgeted Response-Prompt Interaction via Directional Gradient-guided Efficient Token Selection) captures selection-conditioned token utility through a shared validation-directed interaction surrogate valuing each side under the other's retained state. Budgeted alternating selection coordinates retained subsets by aggregating precomputed interactions over the current opposite-side subset to update conditional scores. Structure-aware projection converts conditional response scores into coherent supervision spans. Across three model families, BRIDGE leads compared selection methods overall in mathematical reasoning, code generation, and instruction following. In mathematical reasoning, its advantage over independent selection grows with compression.

CommentsWork in progress

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑