发表机构
Vrije Universiteit Amsterdam(阿姆斯特丹自由大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对实体匹配的LLM方法灵活性不足、忽略推理成本的问题,提出CaRL-EM强化学习控制器,自适应选择算子与模型规模,在7个基准上实现更优质量-成本权衡与零样本迁移。
AI 中文摘要
实体匹配(EM)需要细粒度的上下文理解和领域知识。近期研究表明,大语言模型(LLM)可作为跨领域的强匹配器,但多数方法要么做出独立的成对决策,要么依赖人工设计的复合流水线,在现实的多候选场景中缺乏灵活性,同时通常忽略大规模推理的成本。我们将基于LLM的带候选的实体匹配问题形式化为成本感知的序贯决策问题,并提出CaRL-EM,一种管理LLM操作的强化学习控制器。给定锚定记录的状态、其候选集和成本,CaRL-EM会自适应选择不同的算子(匹配/比较/选择/决策)和模型规模,以最大化质量-成本目标。该策略与抽象算子交互,使得同一控制器在推理时可复用不同的底层LLM后端,无需重新训练。在7个基准数据集上的实验表明,CaRL-EM:(i)学会根据任务复杂度动态规划低成本与高成本算子的使用;(ii)在不同数据集和领域间实现稳健的零样本迁移;(iii)始终比强LLM基线和人工设计的流水线实现更优的质量-成本权衡,在相当或更高质量下推理成本更低。
英文摘要
Entity matching (EM) requires fine-grained contextual understanding and domain knowledge. Recent work shows that large language models (LLMs) can serve as strong matchers across domains, but most methods either make independent pairwise decisions or rely on manually designed composite pipelines, thus lacking flexibility in realistic multi-candidate settings. At the same time, they typically ignore inference cost at scale. We formulate LLM-based EM with candidates as a cost-aware sequential decision problem and propose CaRL-EM, a reinforcement learning controller that manages LLM operations. Given the state of an anchor record, its candidate set, and the cost, CaRL-EM adaptively chooses among different operators (Match/Compare/Select/Decide) and model capacities to maximize a quality-cost objective. The policy interacts with abstract operators, allowing the same controller to be reused with different underlying LLM backends at inference time without retraining. Experiments on 7 benchmarks show that CaRL-EM (i) learns to dynamically plan the usage of inexpensive and expensive operators based on task complexity, (ii) achieves robust zero-shot transfer across diverse datasets and domains, and (iii) consistently achieves a better quality-cost trade-off than strong LLM-based baselines and manually designed pipelines, yielding a lower inference cost at comparable or higher quality.
CommentsAccepted to ACL 2026 Main Conference