arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MERGED:通过生成式专家推理蒸馏实现多模态实体解析

MERGED: Multimodal Entity Resolution via Generated Expert Reasoning Distillation

You-Lin Chen, Kyoungjun Park, Bin Xu, Prithviraj Sen, Abhishek Prasad Ram Tripathi, Pedro Herrero-Vidal

arXiv 2609.01913首次发表:更新:

发表机构

Amazon; Department of Computer Science, The University of Texas at Austin(亚马逊; 德克萨斯大学奥斯汀分校计算机系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MERGED是一种无需人工标注的多模态实体解析蒸馏框架,可将大型VLM的标签与结构化推理迁移至7B参数学生模型,适配新关系定义时仅需10K样本,性能优于人工标注训练的模型及更大规模基准模型,适配速度快、成本低,适合工业部署。

AI 中文摘要

在产品实体解析领域,关系定义会随业务需求不断演变,而传统上适配每一次变更都需要缓慢且成本高昂的人工标注,这类标注通常存在噪声且不包含推理过程。零样本提示的大型视觉语言模型(VLM)能够立即适配新的关系定义,并提供人工标注所缺乏的推理,但它们在生产规模下的成本和延迟是难以承受的。本文提出MERGED,这是一种蒸馏框架,它不仅将标签,还将结构化推理从大型教师VLM迁移到紧凑的7B参数学生模型,整个过程无需人工标注。多位教师为每一对产品打上标签并阐述其决策背后的推理:一致对用于监督微调,而分歧对则由元裁判解决为偏好对,用于直接偏好优化。在多语言电子商务数据集上以人工标注的真值进行评估,得到的学生模型相比用人工标签训练的同一骨干模型,PR-AUC提升了13.79%,并且在成本降低6倍的情况下,比更大规模的Qwen2.5-32B-VL基准模型超出6.32%,同时还实现了更紧密的标签-推理对齐(比Qwen2.5-32B-VL高出10%以上)。此外,从现有检查点重新应用MERGED,仅用10K个样本即可适配新的关系定义,相比零样本设置,PR-AUC提升了6.97%,并且优于从头开始训练的模型。MERGED能够快速适配不断演变的关系定义,可在数天而非数月内支持新关系定义的适配,其成本和延迟适合大规模工业部署。

英文摘要

In product entity resolution, relationship definitions constantly evolve with business needs, yet adapting to each change traditionally requires slow, costly human annotation that is often noisy and carries no reasoning. Large vision-language models (VLMs) prompted zero-shot can adapt to a new definition immediately and supply the reasoning that human labels lack, but their cost and latency are prohibitive at production scale. We present MERGED, a distillation framework that transfers not just labels but structured reasoning from large teacher VLMs into a compact 7B-parameter student, requiring no human annotation for training. Multiple teachers label each product pair and articulate the reasoning behind their decision: agreement pairs supply supervised fine-tuning, while disagreements are resolved by a meta-judge into preference pairs for Direct Preference Optimization. Evaluated against human-labeled ground truth on a multilingual e-commerce dataset, the resulting student improves PR-AUC by 13.79% over the same backbone trained on human labels and surpasses the larger Qwen2.5-32B-VL baseline by 6.32% at 6x lower cost, while also yielding tighter label-reasoning consistency (over 10% above Qwen2.5-32B-VL). Moreover, re-applying MERGED from an existing checkpoint adapts to a new relationship definition with only 10K samples, improving PR-AUC by 6.97% over zero-shot and outperforming from-scratch training. MERGED enables rapid adaptation to evolving relationship definitions, supporting a new one in days rather than months, at a cost and latency suitable for large-scale industrial deployment.

CommentsAccepted at the NeurIPS 2026 Workshop on Grounded and Faithful Vision-Language Models for Real-World Deployment (VLM4RWD). 11 pages, 3 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑