发表机构
Nanjing University; Microsoft Research Asia(南京大学; 微软亚洲研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出PixelJev,一种基于小型开放多模态模型的原生图像决策接口,通过统一识别与多选VQA实现结构化选择,实验表明64样本适配大幅提升Pets准确率并迁移至其他任务,但准确率提升不保证概率校准。
AI 中文摘要
视觉软件通常需要对所提供的备选项做出决策,而非生成解释。我们提出了PixelJev,一个原生图像决策接口,它使用小型开放多模态模型,将图像、任务指令和运行时候选集映射为结构化选择及候选条件概率。其初步实现通过现有的语言模型读出机制统一了识别和多选视觉问答,并分别评估了冻结推理、语言侧适配和留出校准等选项。在七项基准评估中,64样本源适配将Pets准确率从60.13%提升至92.40%(跨优化种子),并迁移至自然重采样、新纹理标签及A-OKVQA而无需目标拟合,而冻结推理已能支持两项VQA任务。在Pets和ScienceQA上进行的匹配提示词对照实验将Pets的大幅提升归因于适配,并识别出适配VQA中候选读出带来的较窄的输出有效性收益。专门的DINOv2探针在源识别上仍更强,冻结的4B模型在DTD和ScienceQA上优于适配的2B模型,且准确率提升并不保证校准的目标概率。这些发现为通用视觉决策模型建立了可行起点,并指出了剩余需求:模式鲁棒性、跨系列迁移以及视觉证据的可靠使用。
英文摘要
Visual software often needs a decision over supplied alternatives rather than a generated explanation. We present PixelJev, a native-image decision interface that maps an image, a task instruction, and a runtime candidate set to a structured choice and candidate-conditioned probabilities using small open multimodal models. Its initial realization unifies recognition and multiplechoice visual question answering through an existing language-model readout, with separately evaluated options for frozen inference, language-side adaptation, and held-out calibration. Across seven benchmark evaluations, 64-shot source adaptation raises Pets accuracy from 60.13% to 92.40% across optimization seeds and transfers to natural resampling, new texture labels, and A-OKVQA without target fitting, while frozen inference already supports both VQA tasks. A matched prompt-only follow-up on Pets and ScienceQA attributes the large Pets gain to adaptation and identifies a narrower output validity benefit of candidate readout in adapted VQA. Specialist DINOv2 probes remain stronger on source recognition, frozen 4B is stronger than adapted 2B on DTD and ScienceQA, and accuracy gains do not ensure calibrated target probabilities. These findings establish a working starting point for general-purpose visual decision models and identify the remaining requirements: schema robustness, cross-family transfer, and reliable use of visual evidence.