arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 4797 信号源:cs.CL, cs.AI, cs.LG

1. 长上下文与记忆 4797 篇

2605.07234 2026-05-11 cs.CL cs.AI 88%

Reformulating KV Cache Eviction Problem for Long-Context LLM Inference

为长上下文LLM推理重构KV缓存淘汰问题

Tho Mai, Joo-Young Kim

机构 * KAIST(韩国科学技术院)

专题命中 长上下文与记忆 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出一种基于输出的分层矩阵乘法近似方法,通过LaProx策略量化token贡献并考虑跨头依赖,实验表明在仅使用5%的KV缓存时保持性能并优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06676 2026-05-11 cs.LG cs.CL 88%

LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction

LKV:端到端学习头部预算和令牌选择以LLM KV缓存淘汰

Enshuai Zhou, Yifan Hao, Chao Wang, Rui Zhang, Di Huang, Jiaming Guo, Xing Hu, Zidong Du, Qi Guo, Yunji Chen

机构 * University of Science and Technology of China(中国科学技术大学) State Key Lab of Processors, Institute of Computing Technology, CAS, Beijing, China(中国科学院计算技术研究所过程器重点实验室,北京,中国) University of Chinese Academy of Sciences, Beijing, China(中国科学院大学,北京,中国)

专题命中 长上下文与记忆 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 本文提出LKV,通过端到端可微优化问题实现KV缓存压缩,学习任务优化全局预算和内在KV重要性,提升长上下文推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13684 2026-04-21 cs.CL cs.AI 88%

HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference

HeteroCache:一种用于长上下文LLM推理的异构KV缓存压缩动态检索方法

Zhiyuan Shi, Qibo Qiu, Feng Xue, Zhonglin Jiang, Li Yu, Jian Jiang, Xiaofei He, Wenxiao Wang

机构 * School of Software Technology, Zhejiang University(浙江大学软件学院) China Mobile (Zhejiang) Research & Innovation Institute(中国移动(浙江)研究院) State Key Lab of CAD&CG, Zhejiang University(浙江大学CAD与CG国家重点实验室) Geely Automobile Research Institute (Ningbo) Co., Ltd.(吉利汽车研究院(宁波)有限公司) FABU Inc.(FABU公司)

专题命中 长上下文与记忆 :LLM(title,title_cn);分类 cs.CL、cs.AI

AI总结 本文提出HeteroCache,通过动态检索方法解决KV缓存压缩问题,利用头部异质性和空间冗余性,实现高效缓存分配和存储机制,提升长上下文推理性能。

Comments Accepted to ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05262 2026-01-12 cs.IR cs.AI cs.CL 88%

LLM2IR: simple unsupervised contrastive learning makes long-context LLM great retriever

LLM2IR: 简单的无监督对比学习使长上下文LLM成为强大的检索器

Xiaocong Yang

机构 * Computer Science(计算机科学)

专题命中 长上下文与记忆 :LLM(title,abstract);large language model(abstract);language model(abstract);pretraining(abstract)

AI总结 LLM2IR通过简单无监督对比学习将长上下文LLM转化为高效检索器,验证了上下文长度与检索能力的关系。

Comments MS Thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05381 2025-10-08 cs.CL cs.AI 88%

Context Length Alone Hurts LLM Performance Despite Perfect Retrieval

Yufeng Du, Minyang Tian, Srikanth Ronanki, Subendhu Rongali, Sravan Bodapati, Aram Galstyan, Azton Wells, Roy Schwartz, Eliu A Huerta, Hao Peng

机构 * University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Amazon.com Inc.(亚马逊公司) USC Information Sciences Institute(南加州大学信息科学研究所) Argonne National Laboratory(阿贡国家实验室) The Hebrew University of Jerusalem(耶路撒冷希伯来大学) University of Chicago(芝加哥大学)

专题命中 长上下文与记忆 :LLM(title,abstract);large language model(abstract);language model(abstract);prompting(abstract)

Comments 18 pages (9 pages of main content), 5 figures, accepted at the Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15253 2025-09-29 cs.CL cs.AI 88%

Conflict-Aware Soft Prompting for Retrieval-Augmented Generation

Eunseong Choi, June Park, Hyeri Lee, Jongwuk Lee

专题命中 长上下文与记忆 :prompting(title,abstract);LLM(abstract);large language model(abstract);language model(abstract)

Comments Accepted to EMNLP 2025; 15 pages; 5 figures, 11 tables; Code available at https://github.com/eunseongc/CARE

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19436 2025-05-27 cs.AI cs.CL 88%

Task Memory Engine: Spatial Memory for Robust Multi-Step LLM Agents

Ye Ye

机构 * New York University(纽约大学)

专题命中 长上下文与记忆 :LLM(title,abstract);large language model(abstract);language model(abstract);prompting(abstract)

Comments Under review. 9 pages main content, 15 pages appendix, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15592 2025-02-24 cs.CL cs.AI 88%

Generalizing From Short to Long: Effective Data Synthesis for Long-Context Instruction Tuning

Wenhao Zhu, Pinzhen Chen, Hanxu Hu, Shujian Huang, Fei Yuan, Jiajun Chen, Alexandra Birch

专题命中 长上下文与记忆 :instruction tuning(title,abstract);large language model(abstract);language model(abstract);post-training(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13009 2024-10-11 cs.CL cs.AI 88%

MetaReflection: Learning Instructions for Language Agents using Past Reflections

Priyanshu Gupta, Shashank Kirtania, Ananya Singha, Sumit Gulwani, Arjun Radhakrishna, Sherry Shi, Gustavo Soares

专题命中 长上下文与记忆 :language agent(title,abstract);LLM(abstract);large language model(abstract);language model(abstract)

Comments We release our experimental code at: https://aka.ms/metareflection-code

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14672 2024-10-07 cs.CL cs.AI 88%

Middleware for LLMs: Tools Are Instrumental for Language Agents in Complex Environments

Yu Gu, Yiheng Shu, Hao Yu, Xiao Liu, Yuxiao Dong, Jie Tang, Jayanth Srinivasa, Hugo Latapie, Yu Su

专题命中 长上下文与记忆 :language agent(title,abstract);LLM(abstract);large language model(abstract);language model(abstract)

Comments EMNLP'2024; 18 pages, 8 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25344 2026-08-04 cs.CL cs.AI cs.LG quant-ph 版本更新 88%

A Hamiltonian-Inspired Local-Operator Ansatz for Slimming Large Language Models

一种用于高效大语言模型的通用张量结构压缩方案

Ying Lu, Peng-Fei Zhou, Qi-Xuan Fang, Pan Zhang, Shi-Ju Ran, Gang Su

机构 * School of Physical Sciences, University of Chinese Academy of Sciences(中国科学院大学物理科学学院) Kavli Institute for Theoretical Sciences, University of Chinese Academy of Sciences(中国科学院大学理论科学研究院) Center for Quantum Physics and Intelligent Sciences, Department of Physics, Capital Normal University(首都师范大学量子物理与智能科学中心) Institute of Theoretical Physics, Chinese Academy of Sciences(中国科学院理论物理研究所)

专题命中 长上下文与记忆 :large language model(title);language model(title);LLM(abstract_cn);分类 cs.CL、cs.AI、cs.LG

AI总结 提出张量混合(MixT)方案,通过将密集线性层替换为张量算子混合体,在保持MMLU准确率的同时大幅减少参数、FLOPs和内存。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22161 2026-07-27 cs.SE cs.CR 新提交 88%

HarnessLLM: Rust Verification Harness Generation with Large Language Models

HarnessLLM:使用大语言模型生成Rust验证框架

Minghua Wang, Yuwei Liu, Lin Huang

专题命中 长上下文与记忆 :large language model(title,abstract);language model(title,abstract)

AI总结 研究针对Rust代码内存安全验证开发验证框架难的问题,提出HarnessLLM自动化工作流程,利用大语言模型从测试套件生成框架,经实验评估效果良好,能检测内存安全漏洞,是首个用大语言模型为Rust项目内存安全验证生成框架的工作。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27366 2026-07-07 cs.AI cs.CL cs.LG cs.MA 版本更新 88%

MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation

MUSE-Autoskill: 通过技能创建、记忆、管理和评估实现自我进化智能体

Huawei Lin, Peng Li, Jie Song, Fuxin Jiang, Tieying Zhang

机构 * ByteDance Inc.(字节跳动公司) Rochester Institute of Technology(罗切斯特理工学院)

专题命中 长上下文与记忆 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 提出MUSE-Autoskill框架,通过统一的技能生命周期(创建、记忆、管理、评估和优化)使LLM智能体持续提升任务解决能力,实验表明生命周期管理的技能可提高任务成功率、效率、复用性和跨智能体迁移。

Comments 30 pages, 9 figures, 15 tables, Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24187 2026-06-25 cs.CV 新提交 88%

Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling

面向多模态大语言模型的长视频快速有效理解:自适应准高斯采样

Kun Zhang, Chenxin Fang, Tao Chen, Baiyang Song, Yunhang Shen, Yiyi Zhou, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(厦门大学多媒体可信感知与高效计算教育部重点实验室)

专题命中 长上下文与记忆 :large language model(title,abstract);language model(title,abstract)

AI总结 提出自适应无训练帧采样方法AdaQ,基于高斯分布3-σ规则动态调整采样区间,在仅用64帧下使Qwen3-VL-8B平均超越GPT4o 15.8%,显著提升长视频理解的鲁棒性和效率。

Comments NeurIPS 2026 submission. 15 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12236 2026-06-12 cs.RO cs.CV 新提交 88%

DrivingAgent: Design and Scheduling Agents for Autonomous Driving Systems

DrivingAgent: 自动驾驶系统的设计与调度智能体

Zhongyu Xia, Wenhao Chen, Yongtao Wang, Ming-Hsuan Yang

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王选计算机技术研究所) University of California, Merced(加州大学默塞德分校)

专题命中 长上下文与记忆 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);foundation model(abstract)

AI总结 提出DrivingAgent框架,通过自动化模块开发(设计阶段)和强化学习训练的轻量级LLM实时调度(调度阶段),解决自动驾驶系统集成新模型和满足实时约束的挑战,在nuScenes和Bench2Drive上取得更优速度-精度权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02474 2026-05-26 cs.CL cs.AI cs.LG 88%

MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents

MemSkill:面向自进化智能体的可学习与进化记忆技能

Haozhen Zhang, Quanyu Long, Jianzhu Bao, Tao Feng, Weizhi Zhang, Haodong Yue, Wenya Wang

机构 * Nanyang Technological University(南洋理工大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Illinois Chicago(伊利诺伊大学芝加哥分校) Tsinghua University(清华大学)

专题命中 长上下文与记忆 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 提出MemSkill框架,将记忆操作转化为可学习和可进化的技能,通过控制器选择技能、执行器生成记忆、设计者进化技能集,形成闭环提升LLM智能体任务性能。

Comments Code is available at https://github.com/ViktorAxelsen/MemSkill

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16538 2026-03-09 cs.CV 88%

OnlineSI: Taming Large Language Model for Online 3D Understanding and Grounding

OnlineSI: 通过大规模语言模型实现在线3D理解和定位

Zixian Liu, Zhaoxi Chen, Liang Pan, Ziwei Liu

机构 * Tsinghua University(清华大学) Nanyang Technological University(南洋理工大学) Shanghai AI Lab(上海人工智能实验室)

专题命中 长上下文与记忆 :large language model(title,abstract);language model(title,abstract)

AI总结 OnlineSI通过整合3D点云与语义信息,提升大规模语言模型在动态环境中的空间理解和物体识别能力。

Comments Project Page: https://onlinesi.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.13126 2026-02-26 cs.CY 88%

Generative agents in the streets: Exploring the use of Large Language Models (LLMs) in collecting urban perceptions

街道中的生成代理:探索大型语言模型(LLMs)在收集城市感知中的应用

Deepank Verma, Olaf Mumm, Vanessa Miriam Carlow

专题命中 长上下文与记忆 :large language model(title,abstract);language model(title,abstract)

AI总结 本研究利用生成代理探索大型语言模型在模拟城市环境中人类行为中的应用,通过街景图像交互和感知评估提升AI在城市感知中的能力。

Comments 30 Pages, 15 Figures, Submitted in a Journal for Peer review

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.00071 2026-02-10 cs.CL cs.AI cs.LG 88%

YaRN: Efficient Context Window Extension of Large Language Models

YaRN: 大语言模型高效上下文窗口扩展

Bowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico Shippole

机构 * EleutherAI University of Geneva(日内瓦大学)

专题命中 长上下文与记忆 :language model(title,abstract);large language model(title);分类 cs.CL、cs.AI、cs.LG

AI总结 YaRN通过高效方法扩展大语言模型的上下文窗口,实现更长上下文的利用与扩展,超越现有最先进水平。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18515 2025-11-12 cs.MA 88%

Socialized Learning and Emergent Behaviors in Multi-Agent Systems based on Multimodal Large Language Models

Sureyya Akin, Shruti T. Tiwari, Ram Bhattacharya, Sagar A. Raman, Kiran Mohanty, Sita Krishnan

专题命中 长上下文与记忆 :large language model(title,abstract);language model(title,abstract)

Comments We have identified critical issues in the code implementation that severely deviate from Algorithm 1, invalidating all experimental results and conclusions. Despite exhaustive efforts to correct these issues, we find they fundamentally undermine the paper's core claims. To uphold academic integrity and prevent misinformation, we are withdrawing this manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21449 2025-10-27 cs.CV 88%

MoniTor: Exploiting Large Language Models with Instruction for Online Video Anomaly Detection

Shengtian Yang, Yue Feng, Yingshi Liu, Jingrou Zhang, Jie Qin

机构 * College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(人工智能学院,南京航空航天大学) Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education, China(脑机智能技术重点实验室,教育部,中国)

专题命中 长上下文与记忆 :large language model(title,abstract);language model(title,abstract)

Comments Accepted to NeurIPS 2025. The first two authors hold equal contributions

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16257 2025-09-30 cs.CV 88%

Plug-and-Play 1.x-Bit KV Cache Quantization for Video Large Language Models

Keda Tao, Haoxuan You, Yang Sui, Can Qin, Huan Wang

机构 * Zhejiang University(浙江大学) Westlake University(西湖大学) Columbia University(哥伦比亚大学) Rice University(Rice大学) Salesforce AI Research(Salesforce人工智能研究)

专题命中 长上下文与记忆 :large language model(title,abstract);language model(title,abstract)

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14326 2025-09-05 cs.SE 88%

Assessing Large Language Models in Comprehending and Verifying Concurrent Programs across Memory Models

Ridhi Jain, Rahul Purandare

专题命中 长上下文与记忆 :large language model(title,abstract);language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18478 2025-07-25 cs.CR 88%

Scout: Leveraging Large Language Models for Rapid Digital Evidence Discovery

Shariq Murtuza

专题命中 长上下文与记忆 :large language model(title,abstract);language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08248 2025-06-10 cs.CL cs.AI cs.IR cs.LG 88%

Eliciting In-context Retrieval and Reasoning for Long-context Large Language Models

Yifu Qiu, Varun Embar, Yizhe Zhang, Navdeep Jaitly, Shay B. Cohen, Benjamin Han

机构 * University of Edinburgh(爱丁堡大学) Apple(苹果公司)

专题命中 长上下文与记忆 :language model(title,abstract);large language model(title);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01074 2025-05-05 eess.SP 88%

WirelessAgent: Large Language Model Agents for Intelligent Wireless Networks

Jingwen Tong, Wei Guo, Jiawei Shao, Qiong Wu, Zijian Li, Zehong Lin, Jun Zhang

专题命中 长上下文与记忆 :large language model(title,abstract);language model(title,abstract)

Comments This manuscript is an extended version of a previous magazine version and is now submitted to a journal for possible publication. arXiv admin note: text overlap with arXiv:2409.07964

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01553 2025-04-08 cs.DB 88%

Bhakti: A Lightweight Vector Database Management System for Endowing Large Language Models with Semantic Search Capabilities and Memory

Zihao Wu

专题命中 长上下文与记忆 :large language model(title,abstract);language model(title,abstract)

Comments 17 pages,5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18434 2025-03-25 cs.CV 88%

A Simple yet Effective Layout Token in Large Language Models for Document Understanding

Zhaoqing Zhu, Chuwei Luo, Zirui Shao, Feiyu Gao, Hangdi Xing, Qi Zheng, Ji Zhang

专题命中 长上下文与记忆 :large language model(title,abstract);language model(title,abstract)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07463 2025-03-11 cs.HC 88%

GenAIReading: Augmenting Human Cognition with Interactive Digital Textbooks Using Large Language Models and Image Generation Models

Ryugo Morita, Ko Watanabe, Jinjia Zhou, Andreas Dengel, Shoya Ishimaru

专题命中 长上下文与记忆 :large language model(title,abstract);language model(title,abstract)

Comments Accepted at AHs2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06680 2024-11-12 cs.SE 88%

Anchor Attention, Small Cache: Code Generation with Large Language Models

Xiangyu Zhang, Yu Zhou, Guang Yang, Harald C. Gall, Taolue Chen

专题命中 长上下文与记忆 :large language model(title,abstract);language model(title,abstract)

Comments 14 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏