统一多模态模型中的视觉弃权(不执行)
Visual Abstention in Unified Multimodal Models
浏览论文内容
中文总结 AI 辅助
针对统一多模态模型忽视任务规则而盲目生成的问题,提出视觉弃权(不执行)形式化定义及DoD基准,并开发VisTA训练方法,使模型先判断可行性再生成,显著提升不可行请求拒绝率且不损编辑性能。
中文摘要 AI 辅助
统一多模态模型(UMMs)整合了理解与生成能力,然而其生成行为很少受其对任务理解的控制。我们形式化了视觉弃权(不执行)问题:当请求的视觉变换在任务规则下不可能实现时,模型应认识到不存在有效解,明确说明这一点,并拒绝生成。我们引入了Draw-or-Decline(DoD)基准,包含7个任务类别中的1,050对可行-不可行请求,联合衡量编辑成功率和拒绝不可行请求的能力。评估8个UMMs后,我们发现编辑能力与弃权(不执行)是两种不同的能力:即使最强的编辑器,其编辑准确率为68.4%,在普通指令下对不可行请求的拒绝率仅为0.4%。其推理过程揭示了原因:模型很少注意到冲突,反而像请求可行一样规划编辑,常常描述图像中不存在的物体,或悄悄将请求改为自己能完成的任务。显式提示这些UMMs报告不可行性会增加文本拒绝,但会降低编辑准确率。我们提出VisTA(视觉变换与弃权(不执行))训练方法,将可行与不可行示例配对,使模型在决定是否生成之前先判断可行性。我们训练VisTA-BAGEL以执行可行编辑并拒绝不可行请求。在无任何提醒的情况下,它对不可行请求的拒绝率达93.0%,相比最强编辑器的0.4%大幅提升,同时仅错误拒绝0.8%的可行请求。与提醒不同,这不会牺牲编辑准确率:VisTA-BAGEL完成了74.3%的可行编辑,超过8个评估UMMs中的任何一个。
英文摘要
Unified multimodal models (UMMs) integrate understanding and generation, yet their generative behavior is rarely governed by what they understand about the task. We formalize visual abstention: when a requested visual transformation is impossible under the task's rules, the model should recognize that no valid solution exists, state this, and decline to generate. We introduce Draw-or-Decline (DoD), a benchmark of 1,050 feasible-infeasible request pairs across 7 task categories that jointly measures editing success and the refusal of infeasible requests. Evaluating 8 UMMs, we find that editing ability and abstention are distinct capabilities: even the strongest editor, at 68.4% editing accuracy, refuses only 0.4% of infeasible requests under ordinary instructions. Their reasoning shows why: the models rarely notice the conflict, and instead plan the edit as if the request were possible, often describing objects that are not in the image, or quietly change the request into one they can complete. Explicitly prompting these UMMs to report infeasibility increases textual refusals but reduces editing accuracy. We propose VisTA (Visual Transformation and Abstention), a training method that pairs feasible and infeasible examples so that a model judges feasibility before deciding whether to generate. We train VisTA-BAGEL to perform feasible edits and decline infeasible requests. Without any reminder, it refuses 93.0% of infeasible requests, up from 0.4% for the strongest editor, while falsely refusing only 0.8% of feasible ones. Unlike a reminder, this does not cost editing accuracy: VisTA-BAGEL completes 74.3% of feasible edits, more than any of the 8 evaluated UMMs.
发表机构
- University of Southern California(南加州大学)
- University of California San Diego(加利福尼亚大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。