arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

标准条件上下文学习:评估视觉语言模型中的标准转移适应性

Criterion-Conditional In-Context Learning: Evaluating Criterion-Shift Adaptation in Vision-Language Models

Kaiyun Yang, Ruilin Yang, Zhimin Yao, Jikai Wang, Wei Ge

arXiv 2607.02575首次发表:更新:

发表机构

Megvii Technology Inc., Beijing, China(美科科技有限公司,北京,中国)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究提出标准条件上下文学习(CC-ICL)新设置,模型需从上下文推断潜在标准并据此调整预测。为此提出两个评估指标,构建CC-Bench基准。实验发现多数模型有边界偏差,简单多标准训练可改善。

AI 中文摘要

视觉语言模型可通过上下文学习执行新任务,标准设置下决策标准固定。现实中任务标准会变,为此引入CC-ICL设置。提出两个指标评估,构建CC-Bench基准。实验表明多数模型有偏差,简单训练策略可改善。

英文摘要

Vision-language models can perform new tasks without parameter updates through in-context learning (ICL), whose core mechanism is utilizing the support set for task induction. In the standard ICL setting, once the task is induced, its decision criterion remains fixed. However, in real-world applications, many tasks exhibit a stable high-level intent, while their decision criteria shift according to specific requirements. Thus, we introduce a new setting, denoted as Criterion-Conditional In-Context Learning (CC-ICL), where models must infer the latent criterion from context and adjust predictions accordingly under fixed task semantics. To evaluate this capability, we propose two complementary metrics, Criterion Invariance and Criterion Sensitivity, capturing the model's robustness and adaptability under criterion shifts. We further construct CC-Bench, a multi-domain benchmark that supports evaluation under the CC-ICL setting. By employing a dual-level data hierarchy, CC-Bench enables legitimate ground-truth variation conditioned on the active criterion even when the task remains fixed. Experiments on CC-Bench reveal that most models exhibit a rigid boundary bias, struggling to align their decisions with the latent criterion. We also find that even a simple multi-criterion training strategy can significantly reduce this bias, improving Criterion Sensitivity and enabling 7B-scale models to surpass proprietary models without degrading general multimodal performance.

CommentsAccepted by ICML 2026. Code is available at https://github.com/MegviiAlgo-Team/CC-ICL

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑