评估视觉语言模型(VLMs)在亚里士多德式多模态说服任务上的表现
Evaluating VLMs on Multimodal Aristotelian Persuasion Tasks
浏览论文内容
中文总结 AI 辅助
该研究采用ImageArg数据集评估VLMs在亚里士多德式多模态说服任务的表现,发现Qwen系列模型在相关检测任务上性能提升并发布代码。
中文摘要 AI 辅助
视觉语言模型(VLMs)在各类任务中展现出出色性能,但尚未在更复杂任务上得到充分评估。亚里士多德提出的说服模型呈三角形结构,凸显其与个人偏见相关的固有挑战。为评估VLMs在这类复杂任务上的进展,我们采用ImageArg数据集,聚焦Logos、Ethos和Pathos检测任务。研究发现,Qwen系列模型的F1分数有所提升,其中Qwen3在Logos和Pathos任务上表现尤为突出,Qwen2在更复杂的Ethos检测任务上展现出竞争力,我们发布代码以推动该方向的研究。
英文摘要
Vision Language Models (VLMs) have demonstrated exceptional performance across various tasks. However, they have not yet been thoroughly evaluated on more complex tasks. The Persuasion Model, conceived by Aristotle, resembles a triangle shape, which highlights its inherent challenges related to personal biases. To assess the progress of VLMs on these complex tasks, we use the ImageArg datasets, focusing on the Logos, Ethos, and Pathos detection tasks. Our findings indicate that models from the Qwen family achieve improved F1 scores, with Qwen3 performing exceptionally well on the Logos and Pathos tasks, while Qwen2 exhibits competitive performance on the more complex Ethos detection task. We release the code to foster research in this direction.
发表机构
- Saarland University(萨尔大学)
机构由 AI 辅助整理,请以论文原文为准。