arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 4797 信号源:cs.CL, cs.AI, cs.LG

1. 长上下文与记忆 4797 篇

2505.06708 2025-05-13 cs.CL 84%

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Zihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang, Kaiyue Wen, Songlin Yang, Rui Men, Le Yu, Fei Huang, Suozhi Huang, Dayiheng Liu, Jingren Zhou, Junyang Lin

专题命中 长上下文与记忆 :large language model(title);language model(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18205 2025-03-27 cs.CL 84%

Contextually Structured Token Dependency Encoding for Large Language Models

James Blades, Frederick Somerfield, William Langley, Susan Everingham, Maurice Witherington

专题命中 长上下文与记忆 :large language model(title);language model(title);分类 cs.CL

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03102 2025-03-26 cs.CL 84%

Structured Token Retention and Computational Memory Paths in Large Language Models

Jonathan Delena, Augustin Moreau, Dominic Ravensdale, Frederick Chatterton

专题命中 长上下文与记忆 :large language model(title);language model(title);分类 cs.CL

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10494 2025-03-14 cs.CL 84%

Source-primed Multi-turn Conversation Helps Large Language Models Translate Documents

Hanxu Hu, Jannis Vamvas, Rico Sennrich

专题命中 长上下文与记忆 :large language model(title);language model(title);分类 cs.CL

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12641 2024-11-12 cs.CL 84%

DetectBench: Can Large Language Model Detect and Piece Together Implicit Evidence?

Zhouhong Gu, Lin Zhang, Xiaoxuan Zhu, Jiangjie Chen, Wenhao Huang, Yikai Zhang, Shusen Wang, Zheyu Ye, Yan Gao, Hongwei Feng, Yanghua Xiao

专题命中 长上下文与记忆 :large language model(title);language model(title);分类 cs.CL

Comments EMNLP Findings 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05678 2024-06-11 cs.CL 84%

SinkLoRA: Enhanced Efficiency and Chat Capabilities for Long-Context Large Language Models

Hengyu Zhang

专题命中 长上下文与记忆 :large language model(title);language model(title);分类 cs.CL

Comments A rethinking of Short Shifted Attention

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.02757 2024-03-06 cs.CL 84%

In-Memory Learning: A Declarative Learning Framework for Large Language Models

Bo Wang, Tianxiang Sun, Hang Yan, Siyin Wang, Qingyuan Cheng, Xipeng Qiu

专题命中 长上下文与记忆 :large language model(title);language model(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07169 2026-08-10 cs.AI cs.LG 新提交 84%

Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory

智能体记忆蒸馏:用分层教师记忆赋能小型大语言模型智能体

Taeil Kim, Kangsan Kim, Sung Ju Hwang

机构 * KAIST(韩国科学技术院)

专题命中 长上下文与记忆 :LLM(title);language model(abstract);small language model(abstract);分类 cs.AI、cs.LG

AI总结 本研究提出无需训练的Agent Memory Distillation框架,通过从大型教师智能体构建三类分层记忆迁移知识,提升小型大语言模型智能体的工具使用性能,在三个基准上实现显著准确率提升且优于现有基线。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05170 2026-08-07 cs.CL cs.AI 新提交 84%

DREAM: LLM-based Dynamic Role-playing via Event-Aware Memory Graph

DREAM:基于大语言模型的事件感知记忆图动态角色扮演

Zhihao Xiao, Mengting Li, Xintao Wang, Linfeng Li, Limin Shui, Mengqi Ji, Borui Cai

机构 * Hangzhou International Innovation Institute, Beihang University(北京航空航天大学杭州国际创新研究院) School of Computer Science, Fudan University(复旦大学计算机科学技术学院)

专题命中 长上下文与记忆 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 DREAM是一种受ABC认知模型启发的结构化记忆框架,通过EMG提升角色扮演智能体的时间与因果连贯性,在CoSER、LIFECHOICE和TCM上取得最优性能。

Comments Accepted at KDD 2026. Camera-ready version to appear. 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01742 2026-08-06 cs.AI cs.CL 版本更新 84%

MemSIF: From Structured Interactions to Dual-Track Fact Memory for LLM Agents

MemSIF:从结构化交互到面向大语言模型智能体的双轨事实记忆

YuFei Luo, Xiucheng Xu, Zhen Yang

专题命中 长上下文与记忆 :LLM(title,abstract);分类 cs.CL、cs.AI

AI总结 该研究针对大语言模型智能体的长期记忆问题,提出MemSIF框架,通过结构化交互记忆与双轨事实记忆缓解两种错位模式,在两个数据集上的五种主干模型中均取得最高总准确率。

Comments Submitted to AAAI 2027. 19 pages, 10 figures, 18 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01418 2026-08-04 cs.AI cs.LG 新提交 84%

Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning

策略滞后下的回滚复用:面向大语言模型强化学习的前缀归一化策略优化

Wenhao Zhang, Yibo Xie, Rui Wang, Jiahua Yang, Lei Jiang, Zibo Yang, Yawei Wang, Jiali Xu, jasperawang, Haoyang Long, Huan Xiong, alantzhao

机构 * Tencent(腾讯) Harbin Institute of Technology(哈尔滨工业大学) Jinan University(暨南大学) University of Science and Technology of China(中国科学技术大学)

专题命中 长上下文与记忆 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 针对大语言模型强化学习中回滚复用导致的离策略问题,提出前缀归一化策略优化(PNPO),在4个更新周期时其数学推理性能优于GSPO,可降低更新批次需求。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26455 2026-07-30 cs.CL cs.AI 新提交 84%

ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models

ForgetBench:语言模型中长期参数记忆的遗忘动态基准测试

Ruxi Gu, Zhenliang Zhang, Wei Wang

专题命中 长上下文与记忆 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

AI总结 本研究提出ForgetBench基准测试,通过两种评估范式与统一分析框架,揭示现有LLMs在长期知识保留与泛化间难以平衡的问题,为未来模型记忆机制优化提供方向。

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25140 2026-07-29 cs.AI cs.CL cs.GR cs.MA 新提交 84%

How Affect Propagates among LLM Agents: Emergent Emotional Contagion in Crowd Simulation

情感如何在大语言模型智能体之间传播:人群模拟中的涌现情感传染

Funda Durupinar

机构 * University of Massachusetts(马萨诸塞大学)

专题命中 长上下文与记忆 :LLM(title,abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 研究多智能体人群模拟中情感传播,通过感知 - 评估 - 表达循环及借鉴相关模型构建架构,在多场景评估,揭示情感传染动态,包括警报扩散、个性影响等,还发现评估步骤动态依赖后端。

Comments 31 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22389 2026-07-27 cs.AR cs.AI cs.LG 新提交 84%

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding

HiKV:用于大语言模型解码的具有硬件加速的分层重要性感知键值缓存

Chao Fang, Jun Yin, Man Shi, Marian Verhelst

机构 * ESAT-MICAS, KU Leuven(KU莱顿大学ESAT-MICAS)

专题命中 长上下文与记忆 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 针对长上下文大语言模型解码时键值缓存的内存瓶颈,HiKV提出算法-硬件协同设计,通过分层重要性感知在两粒度压缩缓存,开发专用加速器统一两阶段加速,在大语言模型上评估实现加速、降能及减少内存访问,优于现有方法。

Comments To appear in the IEEE Transactions on Circuits and Systems I: Regular Papers (TCAS-I)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11614 2026-07-14 cs.CL cs.AI 新提交 84%

Extending LLM Context via Associative Recurrent Memory

通过关联循环记忆扩展大语言模型上下文

Gleb Kuzmin, Ivan Rodkin, Aydar Bulatov, Yuri Kuratov, Lyudmila Rvanova, Mikhail Katkov, Ilia Sochenkov, Misha Tsodyks, Timothy Baldwin, Mikhail Burtsev, Artem Shelmanov

机构 * FusionBrain Lab(融合大脑实验室) MBZUAI(穆罕默德·本·扎耶德人工智能大学) Cognitive AI Systems Lab(认知人工智能系统实验室) RUDN(俄罗斯人民友谊大学) London Institute for Mathematical Sciences(伦敦数学科学研究所) MIRAI(未来人工智能研究机构) Lomonosov Moscow State University(莫斯科国立罗蒙诺索夫大学) Laboratory for Analysis and Controllable Text Generation Technologies RAS(俄罗斯科学院分析与可控文本生成技术实验室) School of Natural Sciences, Institute for Advanced Study, Princeton(普林斯顿高等研究院自然科学学院) Department of Brain Sciences, Weizmann Institute of Science(魏茨曼科学研究所脑科学系)

专题命中 长上下文与记忆 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 研究如何通过关联循环记忆扩展大语言模型上下文长度。提出ARMT方法,构建特定数据集,给出综合训练方法。实验表明该方法能让模型处理超长输入,更好泛化,且降低计算量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26629 2026-06-26 cs.LG cs.CL 新提交 84%

From Weights to Features: SAE-Guided Activation Regularization for LLM Continual Learning

从权重到特征:SAE引导的激活正则化用于大语言模型持续学习

Evan Ning, Wei Xue, Dong Lou, Yike Guo

机构 * The Hong Kong University of Science and Technology(香港科技大学)

专题命中 长上下文与记忆 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 针对大语言模型持续学习中的灾难性遗忘问题,提出利用预训练稀疏自编码器(SAE)在激活空间进行正则化,通过特征掩码平衡稳定性和可塑性,无需旧任务数据,内存效率高,在TRACE和MedCL基准上超越EWC等方法。

Comments 21 pages, 4 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12703 2026-06-12 cs.CR cs.AI cs.LG 新提交 84%

SMSR: Certified Defence Against Runtime Memory Poisoning in Persistent LLM Agent Systems

SMSR:针对持久化LLM代理系统中运行时内存投毒的认证防御

Tarun Sharma

机构 * Independent Researcher(独立研究者)

专题命中 长上下文与记忆 :LLM(title,title_cn);分类 cs.AI、cs.LG

AI总结 提出SMSR防御框架,通过写入时HMAC签名和查询时随机化内存消融与基于判决的多数投票,首次为多会话内存投毒攻击提供认证鲁棒性保证。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10435 2026-06-10 cs.LG cs.CL 新提交 84%

Parallel Causal Associative Fields: Gated Sparse Memory for Long-Context Language Modeling

并行因果关联域:用于长上下文语言建模的门控稀疏记忆

Muhammad Ahmed

机构 * Independent Researcher(独立研究员)

专题命中 长上下文与记忆 :language model(title,abstract);pretraining(abstract);分类 cs.CL、cs.LG

AI总结 提出并行因果关联域(PCAF),通过哈希桶存储局部记录、检索候选集形成稀疏缓存,并与参数化语言模型门控混合,实现稀疏长上下文访问,避免固定状态瓶颈。

Comments 17 pages, 5 figures, and 6 tables. Experiments on WikiText-103, PG-19, and WikiText-2 using TPU v4-32 and NVIDIA RTX 3060 hardware. Code: https://github.com/ahmed123hds/PCAF

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12851 2026-06-01 cs.CL cs.AI 84%

MeMo: Towards Language Models with Associative Memory Mechanisms

MeMo:迈向具有联想记忆机制的语言模型

Fabio Massimo Zanzotto, Elena Sofia Ruzzetti, Giancarlo A. Xompero, Leonardo Ranaldi, Davide Venditti, Federico Ranaldi, Cristina Giannone, Andrea Favalli, Raniero Romagnoli

机构 * Human-centric ART, University of Rome Tor Vergata(人文导向的ART,罗马大学Tor Vergata) University of Edinburgh(爱丁堡大学) Almawave S.p.A.(Almawave公司)

专题命中 长上下文与记忆 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

AI总结 提出MeMo架构,通过分层联想记忆直接记忆文本,实现透明化和模型编辑,实验证明单层和多层配置的记忆能力。

Journal ref Proceedings of Association for Computational Linguistics (Findings), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25680 2026-05-26 cs.CL cs.AI 84%

Simulating Human Memory with Language Models

用语言模型模拟人类记忆

Qihan Wang, Nicholas Tomlin, Michael Hu, Brian Dillon, Tal Linzen

机构 * NYU(纽约大学) UMass Amherst(马萨诸塞大学阿姆赫斯特分校)

专题命中 长上下文与记忆 :language model(title,abstract);prompting(abstract);分类 cs.CL、cs.AI

AI总结 本研究通过心理学经典记忆实验对比语言模型与人类记忆,发现未经调优的模型记忆优于人类,但通过提示策略和压缩器可使模型遗忘方式更接近人类,从而在下游教育任务中成为更有效的用户模拟器。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21240 2026-05-21 cs.LG cs.AI 84%

APEX: Autonomous Policy Exploration for Self-Evolving LLM Agents

APEX:自主策略探索用于自演化大语言模型代理

Yibo Li, Jiashuo Yang, Zhi Zheng, Zhiyuan Hu, Yuan Sui, Shizun Wang, Yufei He, Bryan Hooi

机构 * National University of Singapore(新加坡国立大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 长上下文与记忆 :LLM(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出APEX,一种用于自演化大语言模型代理的自主策略探索方法,通过构建和维护显式的策略空间来解决探索崩溃问题,并在多个基准测试中表现出色。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00505 2026-05-19 cs.IR cs.AI cs.CL 84%

LLM-Oriented Information Retrieval: A Denoising-First Perspective

面向大语言模型的信息检索:一种去噪优先的视角

Lu Dai, Liang Sun, Fanpu Cao, Ziyang Rao, Cehao Yang, Hao Liu, Hui Xiong

机构 * Hong Kong University of Science and Technology(香港科技大学) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 长上下文与记忆 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种以去噪为核心的信息检索方法,强调在信息检索全流程中,最大化可利用证据密度和可验证性是关键瓶颈,通过四个阶段框架和信号-噪声优化技术分类,探讨了信息检索中的挑战和解决方案。

Comments SIGIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16746 2026-05-19 cs.AI cs.LG 84%

State Contamination in Memory-Augmented LLM Agents

内存增强型大语言模型代理中的状态污染

Yian Wang, Agam Goyal, Yuen Chen, Hari Sundaram

机构 * Department of Computer Science, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机科学系)

专题命中 长上下文与记忆 :LLM(title,abstract);分类 cs.AI、cs.LG

AI总结 研究探讨了内存增强型大语言模型代理中由于状态污染导致的安全问题,通过分析内存总结中的毒性内容传播,提出了一种新的衡量指标,并指出在信息压缩前进行净化可以有效减少潜在影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15343 2026-05-18 cs.AI cs.LG cs.MA 84%

Belief Engine: Configurable and Inspectable Stance Dynamics in Multi-Agent LLM Deliberation

信念引擎:多智能体大语言模型协商中的可配置和可检查立场动态

Joshua C. Yang, Maurice Flechtner, Damian Dailisan, Michiel A. Bakker

机构 * ETH Zurich(苏黎世联邦理工学院) Centre for Democracy Studies Aarau, University of Zurich(苏黎世大学民主研究中心) Massachusetts Institute of Technology(麻省理工学院)

专题命中 长上下文与记忆 :LLM(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出Belief Engine,通过可配置的信念更新机制,研究多智能体协商中的立场动态,揭示立场变化背后的证据吸收与锚定因素。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14177 2026-05-15 cs.IR cs.AI cs.CL 84%

Thinking Ahead: Prospection-Guided Retrieval of Memory with Language Models

前瞻性引导的记忆检索:语言模型中的记忆检索

Harshita Chopra, Krishna Kant Chintalapudi, Suman Nath, Ryen W. White, Chirag Shah

机构 * University of Washington(华盛顿大学) Microsoft Research(微软研究院)

专题命中 长上下文与记忆 :language model(title);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 前瞻性引导的记忆检索通过模拟未来步骤来增强长周期检索,提升用户个性化需求的响应质量。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16909 2026-04-28 cs.CL cs.AI 84%

PRISM: Probing Reasoning, Instruction, and Source Memory in LLM Hallucinations

PRISM:探究语言模型幻觉中的推理、指令和来源记忆

Yuhe Wu, Guangyu Wang, Yuran Chen, Jiatong Zhang, Yutong Zhang, Yujie Chen, Jiaming Shang, Guang Zhang, Zhuang Liu

机构 * HKUST(GZ)(香港科技大学(广州)) NYUSH(纽约大学上海分校) DUFE(东华大学) CUHK(SZ)(香港城市大学(深圳)) CUFE(中国金融学院)

专题命中 长上下文与记忆 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 PRISM通过四个维度分析语言模型幻觉,提供细粒度诊断评估,揭示不同模型在指令遵循、记忆检索和逻辑推理上的权衡。

Comments Accepted by ACL main conference 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23040 2026-03-03 cs.CL cs.AI 84%

Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents

回望以推导前行:用于长上下文LLM代理的可回溯记忆

Yaorui Shi, Yuxin Chen, Siyuan Wang, Sihang Li, Hengxing Cai, Qi Gu, Xiang Wang, An Zhang

机构 * University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学) Shanghai Jiao Tong University(上海交通大学) DP Technology(DP科技) Meituan(美团)

专题命中 长上下文与记忆 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 ReMemR1通过整合记忆检索机制和多级奖励设计,提升长上下文问答任务的推理能力与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16411 2026-03-03 cs.CL cs.LG 84%

When Does Divide and Conquer Work for Long Context LLM? A Noise Decomposition Framework

当分治策略在长上下文LLM中何时有效?一种噪声分解框架

Zhen Xu, Shang Zhu, Jue Wang, Junlin Wang, Ben Athiwaratkun, Chi Wang, James Zou, Ce Zhang

机构 * University of Chicago(芝加哥大学) Together AI Duke University(杜克大学) Google DeepMind(谷歌DeepMind) Stanford University(斯坦福大学)

专题命中 长上下文与记忆 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 本文提出了一种噪声分解框架,分析了长上下文LLM中分治策略的有效条件,揭示了任务噪声、模型噪声和聚合噪声的区分,并通过实验验证了多代理分块策略在处理长上下文任务中的有效性。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04919 2026-02-10 cs.CL cs.AI 84%

BudgetMem: Learning Selective Memory Policies for Cost-Efficient Long-Context Processing in Language Models

BudgetMem: 为语言模型中的低成本长上下文处理学习选择性记忆策略

Chandra Vamsi Krishna Alla, Harish Naidu Gaddam, Manohar Kommi

机构 * University of Texas at Arlington(德克萨斯大学阿灵顿分校)

专题命中 长上下文与记忆 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

AI总结 BudgetMem通过学习选择性记忆策略,在语言模型中实现低成本长上下文处理,相比基线RAG节省72.4%内存且F1分数仅下降1.0%。

Comments 11 pages, 3 figures, 5 tables. Evaluated on 700 QA pairs across multiple document lengths

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02197 2026-02-03 cs.LG cs.AI 84%

Hierarchical Adaptive Eviction for KV Cache Management in Multimodal Language Models

分层自适应淘汰用于多模态语言模型的KV缓存管理

Xindian Ma, Yidi Lu, Peng Zhang, Jing Zhang

机构 * College of Intelligence and Computing, Tianjin University(智能与计算学院,天津大学)

专题命中 长上下文与记忆 :language model(title,abstract);large language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出HAE框架,通过双注意力修剪和动态解码淘汰策略优化多模态语言模型的KV缓存管理,减少内存使用并提升推理效率。

Comments 10 oages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏