arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34383cs.IR

纠正以预测:多模态属性值提取的伪值纠正

Correcting to Predict: Pseudo-Value Correction for Multimodal Attribute Value Extraction

发表机构阿里巴巴国际数字商业集团
查看机构详情
  • Alibaba International Digital Commerce Group(阿里巴巴国际数字商业集团)

机构由 AI 辅助整理,请以论文原文为准。

Junhao Zhang, Feiran Hu, Xiao Hu, Baoliang Cui, Xiaoyi Zeng

首次发表
浏览论文内容

中文总结 AI 辅助

针对多模态属性值提取中隐式属性易混淆的问题,提出C2P框架,将提取视为伪值纠正过程,利用多模态证据训练纠正行为,在公共基准和工业数据上超越基线,并在AliExpress在线测试中验证了效果。

中文摘要 AI 辅助

产品属性值提取(AVE)是电子商务中的一项基础任务,旨在从多模态产品资料(如文本和图像)中识别预定义属性的特定值。尽管多模态大语言模型(MLLMs)在AVE中展现出潜力,但它们在提取需要联合推理视觉和文本线索的隐式属性时面临挑战,常常混淆语义相似的值。然而,现有方法往往无法解决此类歧义,因为正确的值通常依赖于容易被忽略或覆盖的微妙多模态线索。为解决这一挑战,我们提出了“纠正以预测”(C2P)框架,将属性提取视为一个纠正过程。给定一个初始伪值(如检索到的候选值或占位符),模型学习利用多模态证据对其进行纠正。在训练过程中,多样化的伪值帮助模型学习基于证据的纠正行为,而自一致性精炼阶段进一步降低了对伪值扰动的敏感性。在推理时,固定的占位符触发已学习的纠正行为,实现无需在线检索或迭代精炼的高效单次预测。我们在公共基准和一个大规模工业数据集上评估了C2P。离线结果表明,C2P优于强基线,尤其在模糊属性上取得了显著提升。在AliExpress上的在线A/B测试进一步显示了卖家采用率、属性完整性和用户参与度的一致提升,验证了C2P在实际部署中的有效性和效率。

英文摘要

Product attribute value extraction (AVE) is a fundamental task in e-commerce, aiming to identify specific values of predefined attributes from multimodal product profiles such as text and images. While multimodal large language models (MLLMs) have shown promise for AVE, they face challenges in extracting implicit attributes that require joint reasoning over visual and textual cues, often confusing semantically similar values. However, existing methods often fail to resolve such ambiguities because the correct value often depends on subtle multimodal cues that are easy to miss or override. To address this challenge, we propose Correcting to Predict (C2P), a framework that treats attribute extraction as a correction process. Given an initial pseudo-value such as a retrieved candidate or placeholder, the model learns to correct it using multimodal evidence. During training, diverse pseudo-values help the model learn evidence-based correction behavior, and a self-consistency refinement stage further reduces sensitivity to pseudo-value perturbations. At inference, a fixed placeholder triggers the learned correction behavior, enabling efficient single-pass prediction without online retrieval or iterative refinement. We evaluate C2P on a public benchmark and a large-scale industrial dataset. Offline results show that C2P outperforms strong baselines, with notable gains on ambiguous attributes. Online A/B tests on AliExpress further show consistent improvements in seller adoption, attribute completeness, and user engagement, validating C2P's effectiveness and efficiency in real-world deployment.

补充信息

↑