arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Aphanta:面向多模态推理的任务对齐图像编辑中间产物诊断框架

Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning

Hengyuan Xu, Wei Cheng, Yumeng Ji, Xuanyang Zhang, Xianfang Zeng, Gang Yu, Xingjun Ma

arXiv 2608.26993首次发表:更新:

发表机构

StepFun; Fudan University(阶跃星辰; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Aphanta框架,用于诊断MLLM与图像编辑中间产物的任务对齐性,实验显示图像编辑是专用视觉工作空间,该框架可用于评估相关流水线效用。

AI 中文摘要

显式视觉中间产物可帮助多模态大语言模型(MLLM)外化空间证据与更新后的视觉状态,但其效用取决于图像编辑器能否忠实实现所需变换。本文提出Aphanta,一种用于MLLM→图像编辑器→MLLM流水线的自动化任务发现与闭环诊断框架。Aphanta评估三类条件:直接推理、使用编辑器生成中间产物的推理、使用理想化参考中间产物的推理,以区分潜在视觉提升空间与当前编辑器的实际效用。在20个候选任务及多组编辑器-MLLM组合上,研究发现效用具有强任务依赖性:提升集中于视觉线索注入、接地、反事实状态实现,而需要符号敏感构建或结构外推的中间产物可靠性显著更低。在选定的正任务子集上,本文整合的Qwen流水线将平均任务得分从0.343提升至0.445(+10.2个百分点,相对提升29.7%),同时完整研究保留过滤及未成功任务以暴露边界。这些结果表明图像编辑是专用视觉工作空间而非通用推理机制,并确立Aphanta作为测量任务-表征对齐、编辑器实现及下游流水线效用的可复用协议。

英文摘要

Explicit visual intermediates can help multimodal large language models (MLLMs) externalize spatial evidence and updated visual states, but their utility depends on whether an image editor can faithfully realize the required transformation. We introduce \textbf{Aphanta}, an automated task-discovery and closed-loop diagnostic framework for the MLLM -> image editor -> MLLM pipeline. Aphanta evaluates three conditions---direct reasoning, reasoning with an editor-generated intermediate, and reasoning with an idealized reference intermediate---to separate potential visual headroom from the practical utility of current editors. Across 20 candidate tasks and multiple editor--MLLM combinations, we find that utility is strongly task-conditioned. Gains concentrate in visual cue injection, grounding, and counterfactual state realization, whereas intermediates requiring symbol-sensitive construction or structural extrapolation are substantially less reliable. On the selected positive-task subset, our consolidated Qwen pipeline improves the mean task score from 0.343 to 0.445 ($+10.2$ points; $+29.7\%$ relative), while the full study also retains filtered and unsuccessful tasks to expose the boundary. These results position image editing as a specialized visual workspace rather than a universal reasoning mechanism, and establish Aphanta as a reusable protocol for measuring task--representation alignment, editor realization, and downstream pipeline utility.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑