识别Transcoders中用于时间敏感事实回忆的时间特征
Identifying Temporal Features within Transcoders for Time Sensitive Factual Recall
浏览论文内容
中文总结 AI 辅助
本研究通过transcoder电路追踪首次绘制时间回忆的特征级映射,识别三类时间节点,并在多个模型中验证其并行句法-语义交互机制,为时间敏感事实回忆提供可干预的MLP组件。
中文摘要 AI 辅助
大型语言模型(LLMs)常因训练语料库的矛盾性质而遭受时间错位问题。尽管当前的缓解策略依赖于计算成本高昂的微调或上下文密集的检索增强生成(RAG),但控制时间敏感回忆的内部机制仍未得到充分探索。与先前识别注意力头(attention heads)和MLP层等时间组件的研究不同,我们通过transcoder电路追踪(circuit tracing)隔离单个MLP特征,首次提供了时间回忆的特征级映射。我们识别出三类节点(通用时间节点、年份通用节点和时序语义节点),它们在事实回忆过程中相互作用以生成时间过滤器。通过分析Gemma 2 2B、LLaMA 3.2 1B和Qwen3-4B,我们表明这些特征并非遵循简单的线性管道,而是通过跨层的并行且混合的句法-语义交互来表示时间。此外,我们发现了一类现有EAP-IG方法不可见的更高层时间组件,确立了transcoders作为时间敏感事实回忆中时间可解释性的更完整视角。这些发现为时间敏感事实回忆中的潜在定向干预提供了MLP组件。
英文摘要
Large Language Models (LLMs) suffer from temporal misalignment, often due to the contradictory nature of their training corpora. While current mitigation strategies rely on computationally expensive fine-tuning or context-heavy retrieval augmented generation (RAG), the internal mechanisms governing time-sensitive recall remain under-explored. Unlike prior studies that identify temporal components such as attention heads and MLP layers, we provide the first feature-level map of temporal recall by isolating individual MLP features via transcoder circuit tracing. We identify three node categories (common temporal, common to the year, and chrono-semantic) which interact to generate a temporal filter during factual recall. By analysing Gemma 2 2B, LLaMA 3.2 1B, and Qwen3-4B, we show that these features do not follow a simple linear pipeline but represent time through a parallel and mixed syntactic-semantic interplay across layers. We additionally discover a class of higher-layer temporal components invisible to existing EAP-IG methods, establishing transcoders as a more complete lens for temporal interpretability in time-sensitive factual recall. These findings present MLP components for potential targeted interventions in time-sensitive factual recall