arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过证据感知多智能体推理实现忠实的情感图像字幕生成

Towards Faithful Sentimental Image Captioning via Evidence-Aware Multi-Agent Reasoning

Tiecheng Cai, Zexian Yang, Chao Chen, Shanshan Lin, Xiangwen Liao

arXiv 2607.25789首次发表:更新:

AI 中文总结

研究情感图像字幕生成中平衡情感表达与视觉保真度的问题,提出SEA-Cap系统,通过情感证据挖掘器提取线索,利用多智能体协作工作流程,依视觉证据审核完善字幕,有效减轻幻觉并达最优性能。

AI 中文摘要

情感图像字幕生成(SIC)需要在情感表达和视觉保真度之间取得平衡。现有方法在这种权衡上常常遇到困难,由于局部基础不足和缺乏情感验证机制而导致幻觉。为解决这些限制,我们提出了SEA-Cap,一个用于忠实且基于证据的情感图像字幕生成的情感证据感知多智能体系统。SEA-Cap包含一个情感证据挖掘器,提取结构化的局部情感线索,将情感控制从全局属性转移到可验证的对象级证据。利用这些证据,我们的框架编排了一个协作工作流程,其中生成器、幻觉检查器和仲裁器通过共享黑板迭代地完善字幕。通过明确根据挖掘出的视觉证据审核生成的内容,SEA-Cap确保了情感准确性和事实一致性。在两个基准数据集上的大量实验表明,SEA-Cap有效地减轻了幻觉并实现了当前最优性能。

英文摘要

Sentimental Image Captioning (SIC) requires balancing emotional expression with visual fidelity. Existing methods often struggle with this trade-off, leading to hallucinations due to insufficient local grounding and the lack of sentimental verification mechanisms. To address these limitations, we propose SEA-Cap, a Sentiment-Evidence-Aware Multi-Agent System for faithful and evidence-grounded sentimental image captioning. SEA-Cap incorporates a Sentiment Evidence Miner that extracts structured, local affective cues to shift sentiment control from global attributes to verifiable object-level evidence. Leveraging this evidence, our framework orchestrates a collaborative workflow where a Generator, Hallucination Checker, and Arbitrator iteratively refine captions via a shared blackboard. By explicitly auditing generated content against mined visual evidence, SEA-Cap ensures both sentiment accuracy and factual consistency. Extensive experiments on two benchmark datasets demonstrate that SEA-Cap effectively mitigates hallucinations and achieves state-of-the-art performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑