arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-02-09 至 2026-02-09 共收录 18 信号源:cs.CL, cs.AI, cs.LG

1. 知识编辑与模型理解 18 篇

2602.06181 2026-02-09 cs.CL 89%

Uncertainty Drives Social Bias Changes in Quantized Large Language Models

不确定性驱动量化大语言模型中的社会偏见变化

Stanley Z. Hua, Sanae Lotfi, Irene Y. Chen

机构 * UC Berkeley(伯克利大学) Centre for Computational Medicine at The Hospital for Sick Children(儿童医院计算医学中心) UCSF(旧金山大学) Meta Superintelligence Labs(Meta超智能实验室)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);post-training(abstract);分类 cs.CL

AI总结 研究发现量化过程导致大语言模型中社会偏见的翻转,且不确定性与量化强度影响偏见变化,需进行后续评估以确保可靠性。

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06852 2026-02-09 quant-ph cs.AI 88%

The Quantum Sieve Tracer: A Hybrid Framework for Layer-Wise Activation Tracing in Large Language Models

量子筛子追踪器:一种用于大语言模型分层激活追踪的混合框架

Jonathan Pan

机构 * Jonathan Pan

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 量子筛子追踪器通过混合量子-经典方法揭示大语言模型中分层激活机制的差异,发现不同模型层在事实回忆和干扰抑制中的不同作用。

Comments 4 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05183 2026-02-09 cs.LG cs.AI 86%

Data-Centric Interpretability for LLM-based Multi-Agent Reinforcement Learning

面向大语言模型多智能体强化学习的数据导向可解释性

John Yan, Michael Yu, Yuqi Sun, Alexander Duffy, Tyler Marques, Matthew Lyle Olson

机构 * goodstartlabs(Good Start Labs)

专题命中 知识编辑与模型理解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出Meta-Autointerp方法,利用SAEs和LLM总结器分析多智能体强化学习行为,发现细粒度行为及奖励黑客,验证SAE特征的有效性并提升智能体性能。

Comments authors 1, 2 and 3 contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14478 2026-02-09 cs.CL cs.LG 86%

Estimating Semantic Alphabet Size for LLM Uncertainty Quantification

估计语义字母表大小以用于LLM不确定性量化

Lucas H. McCabe, Rimon Melamed, Thomas Hartvigsen, H. Howie Huang

机构 * George Washington University(乔治·华盛顿大学) LMI Consulting(LMI咨询公司) University of Virginia(弗吉尼亚大学)

专题命中 知识编辑与模型理解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 本文提出了一种改进的语义字母表大小估计器,用于更准确地估计LLM的不确定性,同时保持高可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12021 2026-02-09 cs.SE 85%

LitterBox+: An Extensible Framework for LLM-enhanced Scratch Static Code Analysis

LitterBox+: 一种可扩展的LLM增强Scratch静态代码分析框架

Benedikt Fein, Florian Obermüller, Gordon Fraser

专题命中 知识编辑与模型理解 :LLM(title,abstract);large language model(abstract);language model(abstract)

AI总结 LitterBox+通过将Scratch块式代码转换为文本形式,利用LLM生成能力提升静态代码分析,提供API和界面扩展,支持代码查询与修复,提升学习者编程体验。

Comments ASE 2025 Tool Demonstration Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04918 2026-02-09 cs.LG cs.CL cs.CY 84%

Simulated Adoption: Decoupling Magnitude and Direction in LLM In-Context Conflict Resolution

模拟采纳:在大语言模型上下文冲突解决中解耦幅度与方向

Long Zhang, Fangwei Lin

专题命中 知识编辑与模型理解 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 研究揭示大语言模型在解决上下文冲突时通过几何位移机制实现模拟采纳,而非单纯抑制知识幅度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10273 2026-02-09 cs.CV cs.AI 79%

Probing Perceptual Constancy in Large Vision-Language Models

探测大型视觉-语言模型中的知觉恒常性

Haoran Sun, Bingyang Wang, Suyang Yu, Yijiang Li, Qingying Gao, Haiyun Lyu, Lianyu Huang, Zelong Hong, Jiahui Ge, Qianli Ma, Hang He, Yifan Zhou, Lingzi Guo, Lantao Mei, Maijunxian Wang, Dezhi Luo, Hokin Deng

机构 * Johns Hopkins University(约翰霍普金斯大学) Emory University(埃默里大学) University of Washington(华盛顿大学) University of California San Diego(加州大学圣地亚哥分校) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) University of Southern California(南加州大学) Washington University in St. Louis(圣路易斯华盛顿大学) Shanghai Jiao Tong University(上海交通大学) East China Normal University(华东师范大学) Stanford University(斯坦福大学) University of California, Berkeley(加州大学伯克利分校) University of Michigan(密歇根大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 本文研究了大型视觉-语言模型在颜色、大小和形状恒常性任务中的表现,发现模型在不同任务上的性能存在显著差异。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06964 2026-02-09 cs.LG cs.AI cs.CL 78%

Learning a Generative Meta-Model of LLM Activations

学习大语言模型激活的生成元模型

Grace Luo, Jiahai Feng, Trevor Darrell, Alec Radford, Jacob Steinhardt

专题命中 知识编辑与模型理解 :LLM(title);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究通过训练扩散模型学习大语言模型激活的分布,提出生成元模型以提升干预保真度和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06213 2026-02-09 eess.AS 78%

From Hallucination to Articulation: Language Model-Driven Losses for Ultra Low-Bitrate Neural Speech Coding

从幻觉到明确:基于语言模型的损失函数用于超低比特率神经语音编码

Jayeon Yi, Minje Kim

专题命中 知识编辑与模型理解 :language model(title,abstract)

AI总结 本文提出基于语言模型的损失函数,用于超低比特率神经语音编码,以缓解音素幻觉问题,提升语义一致性与输出质量。

Comments To appear in ICASSP 2026. Demo wavs, code, and checkpoints (currently) availble at https://github.com/stet-stet/lmloss-icassp2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06818 2026-02-09 cs.AI 77%

Wild Guesses and Mild Guesses in Active Concept Learning

主动概念学习中的大胆猜测与温和猜测

Anirudh Chari, Neil Pattanaik

机构 * Massachusetts Institute of Technology(麻省理工学院) University of California, Berkeley(加州大学伯克利分校)

专题命中 知识编辑与模型理解 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文研究了主动概念学习中EIG与PTS策略的权衡,发现EIG在复杂规则中有效但对简单规则表现欠佳,而PTS通过安全查询维持提案有效性,提升收敛速度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26829 2026-02-09 cs.LG cs.CR 77%

Layer of Truth: Probing Belief Shifts under Continual Pre-Training Poisoning

真相层:在持续预训练中毒下的信念转变探究

Svetlana Churina, Niranjan Chebrolu, Kokil Jaidka

机构 * Centre for Trusted Internet \& Community, National University of Singapore, Singapore

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);pretraining(abstract);分类 cs.LG

AI总结 本研究揭示了持续预训练中虚假信息如何取代模型内部事实表示,通过实验发现中毒对模型信念的影响及可逆性,强调了对事实完整性进行监控的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06129 2026-02-09 cs.LG cs.AI 76%

Urban Spatio-Temporal Foundation Models for Climate-Resilient Housing: Scaling Diffusion Transformers for Disaster Risk Prediction

城市时空基础模型用于气候韧性住房:扩散变换器的扩展用于灾害风险预测

Olaf Yunus Laitinen Imanov, Derya Umut Kulali, Taner Yilmaz

机构 * Technical University of Denmark(技术大学) Eskisehir Technical University(埃斯基谢普大学) Afyon Kocatepe University(阿夫yon卡奥塔佩大学)

专题命中 知识编辑与模型理解 :foundation model(title);分类 cs.AI、cs.LG

AI总结 本文提出Skjold-DiT模型,通过整合时空城市数据预测建筑气候风险,并结合交通网络结构提升灾害响应能力。

Comments 10 pages, 5 figures. Submitted to IEEE Transactions on Intelligent Vehicles

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15735 2026-02-09 cs.LG 70%

EigenTrack: Spectral Activation Feature Tracking for Hallucination and Out-of-Distribution Detection in LLMs and VLMs

EigenTrack:基于隐藏激活谱几何的特征追踪用于LLMs和VLMs中的幻觉和分布外检测

Davide Ettori, Nastaran Darabi, Sina Tayebati, Ranganath Krishnan, Mahesh Subedar, Omesh Tickoo, Amit Ranjan Trivedi

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) Capital One AI Labs(Capital One AI实验室) Intel Labs(英特尔实验室)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.LG

AI总结 EigenTrack通过分析隐藏激活的谱几何,提供一种实时检测LLMs和VLMs中幻觉和分布外错误的可解释方法。

Comments 5 pages, submitted to ICASSP 2026, September 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06532 2026-02-09 cs.CR 67%

Dependable Artificial Intelligence with Reliability and Security (DAIReS): A Unified Syndrome Decoding Approach for Hallucination and Backdoor Trigger Detection

具备可靠性和安全性的人工智能(DAIReS):一种统一的综合解码方法用于幻觉和后门触发检测

Hema Karnam Surendrababu, Nithin Nagaraj

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract)

AI总结 DAIReS提出基于综合解码的统一方法,用于检测学习系统中的安全和可靠性违规,包括后门攻击和幻觉检测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06948 2026-02-09 cs.AI cs.LG 62%

Agentic Uncertainty Reveals Agentic Overconfidence

代理不确定性揭示代理过度自信

Jean Kaddour, Srijan Patel, Gbètondji Dovonon, Leo Richter, Pasquale Minervini, Matt J. Kusner

机构 * University College London(伦敦大学学院) University of Edinburgh(爱丁堡大学) Mila - Québec AI Institute(魁北克AI研究所)

专题命中 知识编辑与模型理解 :prompting(abstract);分类 cs.AI、cs.LG

AI总结 研究发现AI代理在任务执行前后对自身成功率的预测存在过度自信,且执行前的低信息评估比执行后的回顾更有效,对抗性提示能提高预测准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06843 2026-02-09 cs.CL cs.AI 62%

The Representational Geometry of Number

数字的表征几何

Zhimin Hu, Lanhao Niu, Sashank Varma

机构 * Georgia Tech(佐治亚理工学院) University of Edinburgh(爱丁堡大学)

专题命中 知识编辑与模型理解 :language model(abstract);分类 cs.CL、cs.AI

AI总结 本研究通过语言模型探讨数字表征的几何结构,发现任务特定表征可通过线性映射在不同子空间间转换,揭示了共享结构与功能灵活性的平衡机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06748 2026-02-09 cs.CV cs.AI 57%

Gold Exploration using Representations from a Multispectral Autoencoder

利用多光谱自编码器进行黄金勘探

Argyro Tsandalidou, Konstantinos Dogeas, Eleftheria Tetoula Tsonga, Elisavet Parselia, Georgios Tsimiklis, George Arvanitakis

机构 * Technology Innovation Institute(技术创新研究所) Institute of Communication and Computer Systems(通信与计算机系统研究所) Geonova

专题命中 知识编辑与模型理解 :foundation model(abstract);分类 cs.AI

AI总结 利用多光谱自编码器提取生成性表示,提升卫星影像中黄金勘探的准确率与效率

Comments Presented in Eurips2025, 1st Workshop: Advances in Representation Learning for Earth Observation

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06652 2026-02-09 cs.AI cs.CV 57%

Same Answer, Different Representations: Hidden instability in VLMs

相同答案,不同表示:VLMs中的隐藏不稳定性

Farooq Ahmad Wani, Alessandro Suglia, Rohit Saxena, Aryo Pradipta Gema, Wai-Chung Kwan, Fazl Barez, Maria Sofia Bucarelli, Fabrizio Silvestri, Pasquale Minervini

机构 * Sapienza University of Rome(罗马萨皮恩扎大学) CNRS(法国国家科学研究中心) University of Edinburgh(爱丁堡大学) University of Oxford(牛津大学) i3S(i3S研究所)

专题命中 知识编辑与模型理解 :language model(abstract);分类 cs.AI

AI总结 本研究揭示了视觉语言模型中隐藏的不稳定性,通过引入新的评估框架发现模型在内部表示漂移、鲁棒性与决策边界等方面存在显著问题。

详情

展开后加载摘要…

URL PDF HTML 收藏