arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ProtoSemImage:用于可解释文档分类的、具备可变形行对齐的图像值原型

ProtoSemImage: Image-Valued Prototypes with Deformable Row Alignment for Interpretable Document Classification

Mohammad Zare, Pirooz Shamsinejadbabaki

arXiv 2610.11460首次发表:更新:

发表机构

AriooBarzan Engineering Team and Information Technology(阿里欧巴赞工程团队与信息技术)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出 ProtoSemImage 模型,以图像作为文档分类的原型,通过可变形行对齐实现二维视觉模板匹配,在十类任务中较向量原型模型提升 4.3-11.8 个百分点,兼具可解释性与性能。

AI 中文摘要

分类模型中的原型几乎总是向量,而向量没有可读形式。本文探究当原型为图像时会发生什么,文档为该问题提供了自然形式,因为文档可被渲染为多通道图像,其中每个 token 对应一个像素,因此类别代表可采用与其所代表输入相同的形状和通道语义。ProtoSemImage 用一个或多个视觉原型表示每个类别:这些原型是四通道 HSV 空间中的原型图像,其通道承载命名的语言因素。Skip-Gram 目标函数通过四维瓶颈端到端学习该颜色空间,话语边界行成为可微的类型差异行,分类简化为二维视觉模板匹配:即文档图像与原型库之间的可变形行对齐,遵循动态时间规整的思路。由于匹配是空间模式比较而非线性读出,模型可报告输入偏离其原型的位置及偏离的通道,且生成头可将每个原型解码回文本。该图像表示有效:在十类任务中,它在所有三个配对随机种子上,比使用向量原型的相同模型表现高出 4.3 至 11.8 个百分点;而基于距离的匹配则未达到该效果。一项保持表示固定、仅交换分类器的诊断实验恢复了序列基线,这将 20.6 个百分点的差距定位在匹配环节而非颜色压缩环节;一项构建的基准测试中,一对文档共享词袋仅排列不同,证实了其设计的布局保留特性。我们报告了正反两方面结果,因为对于以可检查性为核心目的的表示而言,失败模式与增益同样具有信息价值。

英文摘要

Prototypes in classification models are almost always vectors, and a vector has no readable form. This paper asks what happens when a prototype is an image. Documents give the question a natural form, because a document can be rendered as a multi-channel image in which every token becomes a pixel, so a class representative can take the same shape and the same channel semantics as the inputs it stands for. ProtoSemImage represents each class by one or more visual archetypes: prototype images in a four-channel HSV space whose channels carry named linguistic factors. A Skip-Gram objective learns that color space end to end through a four-dimensional bottleneck, discourse boundary rows become differentiable typed difference rows, and classification reduces to 2D visual template matching: a deformable row alignment between a document image and the archetype bank, in the spirit of dynamic time warping. Because the match is a spatial pattern comparison rather than a linear readout, the model reports where an input departs from its archetype and along which channel, and a generative head decodes each archetype back into text. The image representation works: it beats an otherwise identical model with vector prototypes in all three paired seeds, by between 4.3 and 11.8 points on a ten-class task. The distance-based matching does not. A diagnostic that keeps the representation fixed and swaps only the classifier recovers the sequence baselines, which locates a 20.6-point shortfall in the matching rather than in the color compression, and a benchmark built so that a pair of documents shares a bag of words and differs only in arrangement confirms the layout-preservation it was designed for. We report both directions, because for a representation whose whole purpose is inspect ability, the failure modes are as informative as the gains.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑