MM-VeriRec:面向可验证智能体多模态推荐系统的失败引导融合方法
MM-VeriRec: Failure-Guided Fusion for Verifiable Agentic Multimodal Recommendation
浏览论文内容
中文总结 AI 辅助
针对智能体多模态推荐中视觉约束被忽视和不可能任务未被拒绝的问题,提出MM-VeriRec协议与失败引导融合方法,通过确定性属性验证和失败诊断路由,在MM-ML 1M和Amazon Reviews上超越基线,实现可验证的可信推荐。
中文摘要 AI 辅助
图像往往承载着文本元数据仅能暗示的推荐约束。一部电影可能需要呈现暗黑风格,一件产品可能需要极简设计,而视觉上不可能实现的请求应当被拒绝。智能体多模态推荐系统必须基于文本-图像证据进行推理,判断视觉证据何时具有决定性作用,并在不存在有效行动时选择弃权(不执行)。我们提出了MM-VeriRec,一种可验证的多模态推荐协议及失败引导融合方法,用于处理隐藏视觉约束、图像-文本不匹配以及不可能任务的弃权(不执行)问题。MM-VeriRec基于真实电影海报和产品图像数据集构建任务,通过确定性视觉属性验证每项推荐,并将失败转化为可操作的标签:文本陷阱跟随、视觉忽视和错误接受。融合不应仅仅是拼接多模态特征,而应诊断哪个模态发生失败并路由至相应的修复机制。在MM-ML 1M和Amazon Reviews数据集上,更强的文本和视觉嵌入虽能提升检索性能,但无法消除这些失败模式,而失败引导融合则能够做到。自适应属性门读取与验证器检查相同的标签,其分数是与验证器对齐的上界,用于测试分类体系是否正确路由至修复机制。更具信息量的是在非对齐门下的迁移表现:独立推导的留一法CLIP检测器仍能达到0.7028和0.6111的视觉基础成功率,均高于VBPR基线和普通融合方法。文本与视觉之间的差距在两个LLM家族中均得到复现,且有效的修复方式因领域而异。MM-VeriRec既是一个基准测试集,也是一个用于可信智能体多模态推荐系统的实用诊断循环。
英文摘要
Images often carry recommendation constraints that text metadata only hints at. A movie may need to look dark, a product may need a minimal style, and a visually impossible request should be rejected. Agentic multimodal recommenders must reason over text-image evidence, decide when visual evidence is decisive, and abstain when no valid action exists. We introduce MM-VeriRec, a verifiable multimodal recommendation protocol and failure-guided fusion method for hidden visual constraints, image-text mismatch, and impossible-task abstention. MM-VeriRec builds tasks from real movie-poster and product-image datasets, verifies each recommendation with deterministic visual attributes, and converts failures into actionable labels: text-trap following, visual ignorance, and false acceptance. Fusion should not merely concatenate modalities, but should diagnose which modality failed and route to the appropriate repair. Across MM-ML 1M and Amazon Reviews datasets, stronger text and vision embeddings improve retrieval but do not remove these failure modes, whereas failure-guided fusion does. The adaptive attribute gate reads the same tags the verifier checks and its scores are verifier-aligned upper bounds testing whether the taxonomy routes to the correct repair. More informative is transfer under a non-aligned gate: an independently derived leave-one-out CLIP detector still reaches 0.7028 and 0.6111 visual-grounded success, above both a VBPR baseline and plain fusion. The text-versus-visual gap reproduces across two LLM families, and the repair that helps differs by domain. MM-VeriRec is both a benchmark and a practical diagnostic loop for trustworthy agentic multimodal recommendation.
发表机构
- Independent Research United States
- Independent Research
机构由 AI 辅助整理,请以论文原文为准。