arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12766cs.CV

PatchGen:学习用于视觉泛化的图像内软预测子集

PatchGen: Learning Soft Intra-Image Predictive Subsets for Visual Generalization

Zhaorui Tan, Weimiao Yu, Xi Yang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出PatchGen模块,通过学习样本自适应的软图像内预测子集掩码,提升分类器在多种偏移设置下的泛化性能,增强对未知类别的泛化能力,且无需文本监督即可与视觉-语言方法表现相当。

中文摘要 AI 辅助

视觉分类器需在数据偏移、目标偏移及其组合下实现泛化,但现有多数方法聚焦于域不变性,未能解决图像内预测充分性问题。本文提出结构假设:每张图像包含一个样本自适应的“神谕”图像内预测子集,足以完成标签预测,其余图像块则为可能与标签相关的非必要补充上下文。理论分析表明,将预测限制于该神谕子集可保留全图像块表示能达到的贝叶斯风险,同时其复杂度界随神谕子集大小收紧。基于此观点,本文提出PatchGen,这是一种无文本的模块,可学习样本依赖的软预测子集掩码,作为未观测神谕子集掩码的任务驱动代理。具体而言,组织病理学可视化显示,PatchGen为与肿瘤一致的区域分配的分数高于某些频繁共现的炎症上下文。在涵盖所有三类偏移设置的自然图像及组织病理学图像基准上开展的大量实验表明,PatchGen在多数评估配置中,相比匹配骨干网络的基线方法提升了平均性能,增强了对未知类别的泛化能力,且在无文本监督的情况下,仍能与视觉-语言方法表现相当。

英文摘要

Visual classifiers are expected to generalize under data shifts, target shifts, and their combinations, yet most existing methods focus on domain invariance while failing to address intra-image predictive sufficiency. We investigate the structural hypothesis that each image contains a sample-adaptive oracle intra-image predictive subset sufficient for label prediction, while the remaining patches form non-essential complementary context that may correlate with the label. The theoretical analysis shows that restricting prediction to this oracle subset preserves the Bayes risk achievable by the full-patch representation while admitting a complexity bound that tightens with the oracle-subset size. Based on this view, we propose PatchGen, a text-free module that learns a sample-dependent soft predictive-subset mask as a task-driven proxy for the unobserved oracle subset mask. Specifically, histopathology visualizations suggest that PatchGen assigns higher scores to tumor-consistent regions than to some frequently co-occurring inflammatory context. Extensive experiments on natural and histopathological image benchmarks spanning all three shift settings show that PatchGen improves average performance over matched-backbone baselines in most evaluated configurations, enhances generalization to unknown classes, and remains competitive with vision-language methods without text supervision.

↑