发表机构
POSTECH; Graduate School of Artificial Intelligence, POSTECH; Department of Computer Science and Engineering, POSTECH(浦项科技大学; 浦项科技大学人工智能研究生院; 浦项科技大学计算机科学与工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出KoNA基准评估视觉语言模型的选择性不遵从,发现现有模型在该场景下表现不佳,经KoNA示例微调后,模型不遵从准确率提升且可区分需不遵从的成分。
AI 中文摘要
视觉语言模型(VLMs)应能对合理请求给出有用回应,同时拒绝遵从错误、不安全、不可行或无法回答的请求。然而现有基准大多在整体查询层面评估不遵从,假设每个请求要么应被遵从要么需被拒绝。实际中,真实查询常包含可回答内容与应拒绝遵从的成分。本文提出KoNA基准,用于评估VLMs在5类场景下的选择性不遵从:错误前提、视觉不可达、普遍未知、任务可行性、安全性。每个任务评估两种能力:查询级不遵从,以及配对单查询与复合查询下的成分级不遵从。对多种VLMs的评估显示,模型常无法恰当拒绝、纠正或弃权(不执行),且当查询需要选择性不遵从时,这类失败更明显。为解决该问题,本文用KoNA中需选择性不遵从的示例,搭配应直接回答的全可回答集对VLMs微调。微调后的模型在不遵从准确率上大幅提升,同时在全可回答任务上基本保持性能。结果表明,微调后的模型可区分可回答成分与需不遵从的成分,并以符合任务要求的方式回应。
英文摘要
Vision-language models (VLMs) are expected to respond helpfully to appropriate requests while withholding compliance with requests that are incorrect, unsafe, infeasible, or unanswerable. However, existing benchmarks predominantly evaluate non-compliance at the level of the query as a whole, assuming that each request either warrants compliance or requires withholding compliance. In practice, real-world queries can contain a mixture of answerable content and components for which compliance should be withheld. In this paper, we introduce KoNA, a benchmark for evaluating selective non-compliance in VLMs across five categories: False Premise, Visual Inaccessibility, Universal Unknown, Task Feasibility, and Safety. Each task evaluates two capabilities: query-level non-compliance and component-level non-compliance under paired single and compound queries. Our evaluation across diverse VLMs shows that models often fail to refuse, correct, or abstain appropriately, and these failures become more pronounced when queries require selective non-compliance. To address this challenge, we fine-tune VLMs using KoNA examples that require selective non-compliance, together with a fully answerable set that should receive direct answers. Our fine-tuned models achieve substantial improvements in non-compliance accuracy while largely maintaining performance on fully answerable tasks. These results suggest that the fine-tuned models can distinguish between answerable components and those requiring non-compliance and respond in a task-appropriate manner.
CommentsEMNLP 2026 Main Conference (43 pages). Code and dataset available at https://github.com/mz-kim/KoNA