arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14761cs.CVcs.AI

从视觉反馈到文本评论:一种用于图像支撑评论辅助的多智能体视觉-语言框架

From Visual Feedback to Textual Reviews: A Multi-Agent Vision-Language Framework for Image-Grounded Review Assistance

Utsav Kumar Nareti, Ayush Bansal, Kumari Priya, Chandranath Adak, Soumi Chattopadhyay, Muhammad Saqib, Saeed Anwar

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出多智能体视觉-语言框架,从用户上传产品图像生成可编辑评论草稿,通过产品定位、情感估计、证据生成和评论合成四角色协作,在Amazon数据集上验证了可行性。

中文摘要 AI 辅助

在电子商务平台中,以用户上传图像和视频形式呈现的视觉反馈正变得越来越普遍,因为它提供了关于产品质量、缺陷、包装状况和实际使用情况的真实证据。然而,仅凭视觉反馈往往缺乏做出明智决策所需的上下文解释和主观意见,同时由于撰写详细评论需要付出努力,许多用户提供的文本反馈有限。为弥合这一差距,我们引入了图像支撑评论辅助这一新任务,旨在从用户上传的产品图像中生成可编辑的评论草稿。与专注于客观视觉描述的传统图像描述不同,所提出的任务需要产品特定理解、情感估计以及在具有挑战性的真实世界条件下的证据驱动评论撰写,这些条件包括图像质量退化、过度放大、目标模糊和产品部分可见性。我们提出了一种多智能体视觉-语言框架,由四个专门角色组成:产品定位、视觉情感估计、视觉证据生成和评论合成。该框架采用显式中间表示,包括产品实体、预测评分和证据摘要,以提高可解释性和视觉支撑性。在Amazon Reviews Electronics数据集的一个精选子集上进行的实验证明了从视觉反馈生成连贯、产品感知和情感感知的评论草稿的可行性。据我们所知,这是第一项将图像支撑评论辅助表述为多智能体视觉-语言推理问题的研究,为电子商务系统中AI辅助评论撰写迈出了切实的一步。

英文摘要

Visual feedback in the form of user-uploaded images and videos is becoming increasingly common in e-commerce platforms because it provides authentic evidence of product quality, defects, packaging conditions, and real-world usage. However, visual feedback alone often lacks the contextual explanations and subjective opinions necessary for informed decision-making, while many users provide limited textual feedback due to the effort required to compose detailed reviews. To bridge this gap, we introduce image-grounded review assistance, a novel task that aims to generate editable review drafts from user-uploaded product images. Unlike conventional image captioning, which focuses on objective visual description, the proposed task requires product-specific understanding, sentiment estimation, and evidence-driven review composition under challenging real-world conditions, including degraded image quality, excessive zoom-in, target ambiguity, and partial product visibility. We propose a multi-agent vision-language framework consisting of four specialised roles: product grounding, visual sentiment estimation, visual evidence generation, and review synthesis. The framework employs explicit intermediate representations, including product entities, predicted ratings, and evidence summaries, to improve interpretability and visual grounding. Experiments on a curated subset of the Amazon Reviews Electronics dataset demonstrate the feasibility of generating coherent, product-aware, and sentiment-aware review drafts from visual feedback. To the best of our knowledge, this is the first study to formulate image-grounded review assistance as a multi-agent vision-language reasoning problem, providing a practical step toward AI-assisted review authoring in e-commerce systems.

发表机构

  • Indian Institute of Technology Patna(印度理工学院巴特那分校)
  • Indian Institute of Technology Indore(印度理工学院印多尔分校)
  • CSIRO(澳大利亚联邦科学与工业研究组织)
  • University of Western Australia(西澳大学)

机构由 AI 辅助整理,请以论文原文为准。

↑