arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33609cs.LG

只需编辑一次:通过局部演示精炼激励大语言模型的上下文学习能力

You Only Edit Once: Incentivizing In-Context Capability of LLMs via Local Demonstration Refinement

  • Department of Automation, Tsinghua University(清华大学自动化系)

机构由 AI 辅助整理,请以论文原文为准。

Jiarong Wen, Qi Wang, Yun Qu, Yixiu Mao, Heming Zou, Haoang Chi, Lizhou Cai, Yiqin Lv, Kaiyu Zhang, Yuhang Jiang, Xiangyang Ji

AI总结:

针对上下文学习中演示集选择成本高的问题,提出局部演示编辑(LDE)方法,通过训练小型模型Jev-LDE执行单次编辑优化演示集,平衡性能与成本,并在多个基准上提升ICL效果且具迁移性。

AI中文摘要:

上下文学习(ICL)对于提升大语言模型(LLMs)的推理性能至关重要。然而,ICL在LLMs中的有效性很大程度上受演示集选择的影响。对这些集合进行穷举搜索是组合性的,且现有的选择器通常依赖相关性或似然性代理来隐式评估ICL质量。使用这些策略对目标LLM进行重复查询可能会产生高昂成本。本工作将选择问题简化为一个受约束的局部搜索问题,并提出了局部演示编辑(LDE)方法。从一个初始检索到的演示集出发,LDE采用单一结构化编辑来探索其周围邻域,同时平衡性能提升与搜索成本。技术上,LDE被简化为一个策略搜索问题,为此我们训练了一个小型LLM,称为Jev-LDE。该模型作为系统1,通过执行诸如\ exttt{Keep}(保留)、\ exttt{Delete}(删除)或\ exttt{Replace}(替换)元素等操作来修改检索到的演示集,所有这些都在具有可验证奖励的强化学习框架内进行。在测试时,Jev-LDE对检索到的演示集执行一次编辑,随后对目标LLM进行一次推理,避免了迭代上下文评分或子集搜索的需要。在标准分类基准上,使用Jev-LDE作为即插即用模块的各种目标LLM持续提升了ICL性能,且Jev-LDE显示出对保留基准和模型的迁移性,无需重新训练。这些发现表明,LDE方法为利用目标LLM的ICL能力提供了一种高效且适应性强的方法。

英文摘要:

In-context learning (ICL) is crucial for boosting the inference performance of large language models (LLMs). However, the effectiveness of ICL in LLMs is greatly influenced by the choice of demonstration sets. Exhaustive searches over these sets are combinatorial, and existing selectors often rely on relevance or likelihood proxies to implicitly assess ICL quality. Making repeated queries to the target LLM with these strategies can incur substantial costs. This work simplifies selection by framing it as a constrained local search problem and presents local demonstration editing (LDE). Starting with an initially retrieved set of demonstrations, LDE employs a single structured edit to explore its surrounding neighborhood while balancing performance gains with search costs. Technically, LDE is reduced to a policy search problem, for which we train a small LLM, referred to as Jev-LDE. This model as the System-1 modifies the retrieved demonstration set by performing actions such as \texttt{Keep}, \texttt{Delete}, or \texttt{Replace} elements, all within a framework of reinforcement learning with verifiable rewards. At test time, Jev-LDE executes a single edit of the retrieved demonstration set, followed by one inference from the target LLM, avoiding the need for iterative context scoring or subset searches. Across standard classification benchmarks, various target LLMs with Jev-LDE as the plug-and-play module consistently improve ICL performance, and Jev-LDE shows transferability to held-out benchmarks and models without retraining. These findings indicate that the LDE approach offers an efficient and adaptable method for harnessing the ICL capabilities of target LLMs.

↑