arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25023cs.AIcs.CV

CVE-SAI:用于风险可控电商搜索的反事实视觉证据引导的选择性属性索引

CVE-SAI: Counterfactual Visual Evidence-Guided Selective Attribute Indexing for Risk-Controlled E-commerce Search

Xiaolong Sun, Qichao Wang, Hangyu Li, Liang Chen

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出CVE-SAI方法,将电商属性推断与索引准入分离,通过反事实视觉证据控制风险,在5%不安全准入预算下,于Amazon Berkeley Objects数据集上实现了更优的属性补全与检索性能。

中文摘要 AI 辅助

多模态商品模型可补全缺失的电商属性,但现有方法仍仅优化属性-答案准确率,未验证视觉支持,将瞬态预测与持久索引准入混为一谈,且未对事实上错误或视觉无支撑的值进行明确风险控制。我们提出反事实视觉证据引导的选择性属性索引(Counterfactual Visual Evidence-Guided Selective Attribute Indexing,CVE-SAI),该方法首先从主图像和属性问题中推断并冻结一个受本体约束的候选对象(不使用目录文本),随后决定该候选对象是否应进入索引。焦点区域失真(Focus-Zone Distortion,FZD)通过受控反事实干预构建特定于属性的视觉依赖代理,证据引导的注意力重分配(Evidence-Guided Attention Redistribution,EGAR)利用该代理优化受本体约束的评分。在证据必要性、证据保留、干扰变换稳定性及候选特定目录文本冲突审计前,规范候选对象已被冻结;目录文本仅可收紧准入条件,不可修改候选对象。独立家族级校准在5%的不安全准入预算下,通过单侧有限样本边界选择一个策略。针对Amazon Berkeley Objects的5个视觉属性开展的实验表明,CVE-SAI提升了属性推断与证据定位能力,在共享风险协议下实现了最高的认证准入覆盖率,且在自动准入系统中产生了最强的可控检索性能与最低的不安全自动诱导暴露。因此,将推断与准入分离可实现视觉支撑的属性补全,以提升检索效果,同时限制持久索引污染。

英文摘要

Multimodal product models can complete missing e-commerce attributes, yet current methods still optimize attribute-answer accuracy without verifying visual support, conflate transient prediction with persistent index admission, and lack explicit risk control over factually incorrect or visually unsupported values. We address these gaps with Counterfactual Visual Evidence-Guided Selective Attribute Indexing (CVE-SAI), which first infers and freezes an ontology-constrained candidate from the primary image and attribute question without catalog text, and then decides whether that candidate should enter the index. Focus-Zone Distortion (FZD) constructs an attribute-specific visual-dependence proxy through a controlled counterfactual intervention, and Evidence-Guided Attention Redistribution (EGAR) uses the proxy to refine ontology-constrained scoring. The canonical candidate is frozen before evidence necessity, evidence retention, nuisance-transformation stability, and candidate-specific catalog-text conflict audits; catalog text can only tighten admission and cannot revise the candidate. Independent family-level calibration selects one policy with a simultaneous one-sided finite-sample bound under a 5% unsafe-admission budget. Experiments on five visual attributes derived from Amazon Berkeley Objects show that CVE-SAI improves attribute inference and evidence localization, achieves the highest certified admission coverage under the shared risk protocol, and yields the strongest controlled retrieval performance with the lowest unsafe auto-induced exposure among automatic-admission systems. Separating inference from admission therefore enables visually supported attribute completion to improve retrieval while limiting persistent index contamination.

发表机构

  • Sun Yat-Sen University(中山大学)
  • Nanyang Technological University(南洋理工大学)
  • Tencent(腾讯)

机构由 AI 辅助整理,请以论文原文为准。

↑