大语言模型上下文学习中数值序列表示的图信号处理视角
A Graph Signal Processing Perspective on Numerical Sequence Representations in LLM In-Context Learning
浏览论文内容
中文总结 AI 辅助
本文从图信号处理视角研究LLM上下文学习中数值序列的表示,发现数值推理的内部特征随上下文长度和输入动态复杂性变化,且在不同模型家族间一致。
中文摘要 AI 辅助
预训练大语言模型(LLMs)在以文本序列化的数值序列上展现出了上下文学习(ICL)能力。现有研究主要通过预测误差等输出层面评估来识别和表征这类数值推理,但人们对数值信息在LLM表示中的组织方式仍知之甚少。为研究这种内部组织,本文采用图信号处理视角:注意力机制在 token 间诱导出加权图,而 token 隐状态则定义为该图节点上的信号。定量图谱诊断与定性 token 图可视化显示,随着上下文长度增加,表示会因输入动态复杂性而更明显分化:更简单的输入产生的注意力诱导 token 图具有更强的全局连通性、更平滑且谱集中的隐状态信号,而更复杂的输入则产生更局部化的图、谱支持更宽且高频能量更高的隐状态信号。这些发现共同表明,数值 ICL 存在与上下文相关的系统性内部特征,且该特征在不同模型家族间具有一致性。
英文摘要
Pretrained large language models (LLMs) have demonstrated in-context learning (ICL) capabilities for numerical inference over sequences serialized as text. Prior work has identified and characterized this form of numerical inference primarily through output-level evaluations such as prediction error. However, how numerical information is organized within LLM representations remains much less understood. To study this internal organization, we adopt a graph signal processing perspective in which attention induces a weighted graph over tokens, while token hidden states define signals on its nodes. Quantitative graph-spectral diagnostics and qualitative token-graph visualizations reveal that representations become more clearly differentiated by input dynamical complexity as context length increases. Simpler inputs produce attention-induced token graphs with stronger global connectivity and smoother, spectrally concentrated hidden-state signals, whereas more complex inputs produce more localized graphs and hidden-state signals with broader spectral support and greater high-frequency energy. Together, these findings point to systematic, context-dependent internal signatures associated with numerical ICL that are conserved across model families.
发表机构
- Cornell University(康奈尔大学)
- Goodfire AI(古德火人工智能公司)
- Imperial College London(伦敦帝国学院)
机构由 AI 辅助整理,请以论文原文为准。