arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24688cs.DBcs.CLcs.LG

超越规模和生成:理解基于语言模型的实体匹配

Beyond Scale and Generation: Understanding Language Model-based Entity Matching

Zeyu Zhang, Xue Li, Iacer Calixto, Paul Groth, Sebastian Schelter

首次发表
浏览论文内容

中文总结 AI 辅助

研究基于语言模型的实体匹配,通过控制因子研究Qwen3家族的多种架构、变体和大小,评估跨数据集可迁移性与计算成本,明确模型变体对双编码器的关键作用,以及不同匹配器架构性能差异因素,推动相关研究与基准设计发展。

中文摘要 AI 辅助

实体匹配用于识别指代同一现实世界实体的记录。语言模型可通过双编码器、交叉编码器和生成匹配器架构来适应此任务。然而,先前研究常将匹配器架构与模型主干、模型变体(反映不同预训练目标)和模型大小的差异混为一谈,难以分离性能提升的来源。我们通过一项涵盖Qwen3家族的三种匹配器架构、三种模型变体、三种模型大小以及九个数据集的控制因子研究来解决此问题,共进行1215次微调运行。我们还评估了跨数据集可迁移性和计算成本。结果表明,模型变体对双编码器至关重要;交叉编码器始终优于双编码器;生成匹配器并非普遍优于交叉编码器,其优势集中在分布转移情况下。此外,更大的模型更依赖捷径学习,不一定表现更好。这些发现阐明了匹配器架构性能差异的因素,推动未来研究和基准设计更好地将架构选择与模型级因素分离,同时明确评估分布转移和跨数据集可迁移性。我们在https://这个网址发布了实验结果、代码、训练脚本和评估数据。

英文摘要

Entity matching identifies records that refer to the same real-world entity. Language models can be adapted to this task through bi-encoder, cross-encoder, and generative matcher architectures. However, prior studies often conflate matcher architecture with differences in model backbone, model variant(reflecting different pretraining objectives), and model size, making it difficult to isolate the sources of performance gains. We address this issue through a controlled factorial study spanning three matcher architectures, three model variants and three model sizes from the Qwen3 family, and nine datasets, totaling 1,215 fine-tuning runs. We also evaluate cross-dataset transferability and computational cost. Our results show that model variant is critical for bi-encoders: embedding-oriented variants provide stronger initialization and more favorable representation geometry predictive of downstream matching performance. Cross-encoders retain a consistent advantage over bi-encoders because they jointly encode record pairs rather than representing each record independently, although larger models partially narrow this gap. Generative matchers do not universally outperform cross-encoders. Instead, their advantages concentrate under distribution shift, including subtle unseen differences in record schemas and cross-dataset transfer. We further find that larger models rely more heavily on shortcut learning and therefore do not necessarily perform better. These findings clarify the factors underlying performance differences across matcher architectures and motivate future research and benchmark designs that better disentangle architectural choices from model-level factors while explicitly evaluating distribution shift and cross-dataset transferability. We release our experimental results, code, training scripts, and evaluation data at https://github.com/Jantory/llm-trained-matcher.

发表机构

  • University of Amsterdam(阿姆斯特丹大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑