arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11859cs.AI

从参数到答案:大语言模型如何检索和使用其内部知识

From Parameters to Answers: How LLMs Retrieve and Use Their Internal Knowledge

  • University of Science and Technology of China(中国科学技术大学)
  • Singapore Management University(新加坡管理大学)
  • The University of Tokyo(东京大学)
  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

Wenkang Wei, Yuan Fang, Renhe Jiang, Hong Cheng, Xingtong Yu

中文总结 AI 辅助

本研究通过逐层干预揭示LLM回答时对查询路由和内部知识的依赖变化,区分了早期可读性、自然强度、因果引导和后期内容依赖,并发现模型间存在差异。

中文摘要 AI 辅助

语言模型在回答问题时,其对查询路由信息和目标知识的依赖是如何变化的?我们通过在问题结尾处的隐藏状态上进行逐层干预来研究这一问题。在Qwen、Llama和Gemma模型中,我们比较了国家-大陆问题与名词、形容词和代码答案,同时保持若干拟合测量值的区分。一个配对条件请求方向描述了在自然单国问题中查询的是哪个国家;一个全局请求方向描述了配对问题中第一国与第二国请求的差异;单独的选择候选用于测试隐藏状态中已有内容之间的控制。对冻结的Qwen自然问题状态进行的诊断性再分析表明,在对其干预开始改变后续拟合知识之前,配对条件方向会增强,且这一因果窗口在答案支持内容仍在形成时开启。三个模型的配对轨迹并不一致:Gemma显示出部分重叠的中层路由-内容特征,而Llama在相同门控下没有持续的路由效应窗口。在配对协议中,对全局请求方向的依赖从较早的固定层集到较晚的层集逐渐减弱,而对拟合内容的依赖则持续存在。一项匹配的Qwen比较显示,配对条件方向保留了后期效应,因此这种操作交接涉及的是全局拟合方向,而非所有请求信息。这些结果区分了早期可读性、自然强度、因果引导和后期内容依赖。

英文摘要

How does a language model's dependence on query-routing information and target knowledge change as it answers a question? We study this question through layerwise interventions on the hidden state at the end of the question. Across Qwen, Llama, and Gemma, we compare country-continent questions with noun, adjective, and code answers while keeping several fitted measurements distinct. A pair-conditioned request direction describes which country is queried in natural single-country questions; a global request direction describes first- versus second-country requests in paired questions; separate selection candidates test control among contents already available in the hidden state. A diagnostic reanalysis of frozen Qwen natural-question states shows that the pair-conditioned direction grows stronger before interventions on it begin to alter later fitted knowledge, with this causal window opening while answer-supporting content is still forming. The paired three-model trajectories are not uniform: Gemma shows a partially overlapping mid-layer routing-content profile, whereas Llama has no sustained routing-effect window under the same gates. In the paired protocol, dependence on the global request direction decreases from fixed earlier to later layer sets while dependence on fitted content persists. A matched Qwen comparison shows that the pair-conditioned direction retains a late effect, so this operational handoff concerns the global fitted direction rather than all request information. These results separate early readability, natural strength, causal steering, and later content dependence.

补充信息

↑