完整字形图像胜过词元嵌入:Transformer的对照研究
One Image, No Tokens: A Controlled Study of Glyph-Based Chinese Language Modeling
浏览论文内容
中文总结 AI 辅助
挑战用离散词元嵌入表示文本,用字符序列光栅化图像经视觉编码器处理替代,构建双分支框架对比,发现基于视觉模型性能优于基线,揭示优势条件及跨脚本差异。
中文摘要 AI 辅助
现代语言模型通常将文本表示为离散词元嵌入序列,本文对此提出挑战,尤其针对中文,用字符序列的单个光栅化图像完全取代基于索引的词元嵌入,由共享ResNet和浅层视觉Transformer组成的视觉编码器处理。构建双分支对照框架,发现基于视觉的模型始终优于基线,揭示优势条件及跨脚本差异。
英文摘要
Modern language models generally represent text as sequences of discrete token embeddings, an assumption deeply rooted in current practice but rarely questioned. We challenge this representation, especially for Chinese, by replacing index-based token embeddings entirely with a single rasterized image of the character sequence, processed by a vision encoder composed of a shared ResNet and a shallow Vision Transformer. To isolate the role of input representation, we construct a dual-branch controlled framework in which both a Vision-based model and an index-based baseline share an identical decoder backbone, training objective, optimizer, and data curriculum. Any performance difference is therefore attributable to the input modality only. Across all tested decoder backbones, the Vision-based model consistently outperforms the baseline, reaching a peak accuracy of 0.429 versus 0.355 for the index-based baseline---a 21\% relative improvement---while converging in about half the number of training epochs. The advantage emerges especially within the first five epochs (under 21\% of total data) and persists under moderate character corruption: the corrupted Vision model matches the \emph{clean} index-based baseline. Ablation studies reveal that the advantage requires both spatially coherent input and a ViT encoder with 2D positional encodings. A cross-script comparison on English shows the advantage does not transfer directly to alphabetic writing systems, suggesting that the uniform visual density and radical structure of Chinese characters are enabling conditions. These findings suggest that transformers are more modality-agnostic than commonly assumed, and that discrete tokenization is not a fundamental requirement for Chinese language modeling.
发表机构
- Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)
机构由 AI 辅助整理,请以论文原文为准。