arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03711cs.CVcs.CLcs.LG

注意力对字母大小写敏感

Attention is Case-Sensitive

Maximilian Dillitzer, Tin Stribor Sohn, Jason J. Corso, Michael Auerbach

首次发表
浏览论文内容

中文总结 AI 辅助

该研究发现LLMs和VLMs存在大小写效应,即文本中目标信息采用特定大小写格式会调节注意力分配,但该效应不一定提升任务准确率,推理模型的思考阶段可缓解此效应,且该效应可部分迁移至VLMs。

中文摘要 AI 辅助

在人类视觉感知中,大写字母是一种自然的显著性线索,能在小写文本中吸引注意力。本文开展了系统性实证表征研究,揭示了大型语言模型(Large Language Models, LLMs)存在类似特性:字母大小写会调节内部注意力分配。通过对13种模型(9种LLMs和4种视觉语言模型(Vision-Language Models, VLMs))、不同分词方案的分析,我们发现,在小写文本语境中,采用交替大小写或全大写格式呈现目标信息,会使注意力集中在这些文本片段上。该效果在文本领域具有普遍性,在所有被评估的非推理模型中均成立。我们将其视为预训练Transformer此前未被充分探索的潜在特性,而非一种规定性方法。研究揭示了注意力-性能的核心分歧:尽管这种“大小写效应”能稳健地改变注意力分配,但其对下游准确率的影响却不容忽视——注意力集中度提升并不必然提高任务准确率,在交替大小写这类高熵语境中甚至会降低准确率。我们还确定了边界条件:推理模型的审慎“思考”阶段会充当语义缓冲,缓解文本中的排版敏感性。将研究扩展至VLMs,我们发现该效应可部分迁移:相同提示侧的大小写会沿两个耦合轴重组跨模态注意力,主要是从图像宏观脱离至文本提示,其次是剩余视觉注意力集中在目标区域。通过将大小写隔离为一种无需模型访问或微调的零样本注意力引导机制,我们为预训练如何内化排版强调提供了新的基础理解。

英文摘要

In human visual perception, uppercase lettering serves as a natural salience cue that captures attention within lowercase text. In this paper, we present a systematic empirical characterization study revealing that Large Language Models (LLMs) exhibit an analogous property: letter casing modulates internal attention allocation. Through analysis across 13 models, nine LLMs and four Vision-Language Models (VLMs), with diverse tokenization schemes, we show that formatting target information in alternating or uppercase against a lowercase context concentrates attention on those textual spans. In text this effect is universal, holding across every evaluated non-reasoning model. We frame it as a previously under-explored latent property of pretrained transformers rather than a prescriptive method. Our investigation reveals a central attention-performance divergence: while this "casing effect" robustly shifts attention, its impact on downstream accuracy is non-trivial, increased concentration does not inherently improve task accuracy and, in high-entropy contexts like alternating case, can degrade it. We further identify a boundary condition: the deliberative "thinking" phase in reasoning models acts as a semantic buffer that mitigates typographic sensitivity in text. Extending the study to VLMs, we find the effect transfers partially: the same prompt-side casing reorganizes cross-modal attention along two coupled axes, predominantly a macroscopic disengagement from the image toward the text prompt, and secondarily a concentration of the residual visual attention on the target region. By isolating casing as a zero-shot mechanism for attention steering that requires no model access or fine-tuning, we provide a new foundational understanding of how pretraining internalizes typographic emphasis.

补充信息

↑