arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10084stat.MLcs.AIcs.CVcs.LG

零样本学习中偏差的统计方法:以手写识别为视角

A statistical approach to bias in zero-shot learning: the lens of handwriting recognition

Clarence Chew, Gim Siang Chia, Sukalpa Chanda, Subhroshekhar Ghosh, Soumendu Sundar Mukherjee

首次发表
浏览论文内容

中文总结 AI 辅助

针对广义零样本学习中的训练类别偏差问题,提出两阶段统计校正方法,结合黑盒特征学习器与蒙特卡洛偏差校正器,在超大词汇手写识别中实现20%以上相对准确率提升,并发现约15维的低维表示。

中文摘要 AI 辅助

广义零样本学习(GZSL)已成为视觉识别系统的重要范式,这些系统必须泛化到训练期间未观察到的类别。传统GZSL技术受限于其仅适用于相对较少数量的此类未见类别,由于其对训练期间观察到的类别存在众所周知的误分类偏差,超越这一规模的可扩展性具有挑战性。在本工作中,我们通过极大词汇量上的零样本手写单词识别的视角来研究GZSL范式。我们提出了一种纠正此偏差的统计方法,该方法将任何经典GZSL特征学习器视为黑盒机制,其内在偏差在于识别典型数据点的训练状态(已见与未见),我们旨在纠正此偏差,类似于分布外推断问题。我们的方法利用简单的两阶段层次架构,第一阶段结合经典GZSL黑盒,第二阶段结合轻量级蒙特卡洛偏差校正器的集成。一旦去偏,测试数据的分类仅在其预测的训练状态限制下通过基于统计的方法(如最近邻、逻辑回归和随机森林)进行。与已有技术相比,我们在未见单词的分类上实现了超过20%的相对准确率提升。一个关键结果是,大规模词汇表上的单词识别适用于更低维的表示(约15维)。我们的方法由数学分析支撑,该分析捕捉了偏差校正统计方法的本质。我们的偏差纠正方法可以与任何经典GZSL学习器作为黑盒以交钥匙方式结合,从而表明该方法在不同领域的各种GZSL实现中具有广泛适用性。

英文摘要

Generalized zero-shot learning (GZSL) has emerged as an important paradigm for visual recognition systems that must generalize to classes that were not observed during training. Traditional GZSL techniques are limited by their applicability to a relatively small number of such unseen classes, scalability beyond which is challenging due to its well-known misclassification bias towards classes observed during training. In this work, we investigate the GZSL paradigm through the lens of zero-shot handwritten word recognition over extremely large vocabularies. We propose a statistical approach to rectifying this bias, which views any classical GZSL feature learner as a black box mechanism whose intrinsic bias in identifying the training status (seen vs. unseen) of a typical data point we aim to correct, similar to an out of distribution inferential problem. Our method leverages a simple two-stage hierarchical architecture, combining a classical GZSL blackbox in the first stage and an ensemble of lightweight Monte Carlo bias-correctors in the second. Once debiased, the classification of test data is undertaken only restricted to its predicted training status via well-founded statistical methods (eg nearest neighbour, logistic regression and random forests). We achieve relative accuracy improvements of over 20% in the classification of unseen words compared to established techniques. A key outcome is that word recognition over large scale vocabularies is amenable to a much lower dimensional representation (~15 dimensions). Our approach is underpinned by mathematical analysis that captures the essence of the statistical approach to bias correction. Our approach to bias rectification can be combined in a turn-key fashion with any classical GZSL learner as a blackbox, thereby suggesting a wide scope of applicability of this method for a wide variety of GZSL implementations in different domains.

发表机构

  • National University of Singapore(新加坡国立大学)
  • Østfold University College(东福尔大学学院)
  • Indian Statistical Institute(印度统计研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑