arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 22368 信号源:cs.CL, cs.AI, cs.LG

1. 效率与部署 22368 篇

2606.08129 2026-06-09 cs.AI 新提交 89%

Cross-LLM Consistency in Inference: Evidence from Shared Interactions

推理中的跨LLM一致性:来自共享交互的证据

Siyu Lou, Yao Yan, Yuntian Chen, Quanshi Zhang

机构 * School of Computer Science Shanghai Jiao Tong University(上海交通大学计算机科学学院) Ningbo Key Laboratory of Advanced Manufacturing Simulation Eastern Institute of Technology, Ningbo(宁波市先进制造仿真重点实验室,宁波东方理工大学) College of Computer and Information Science Chongqing Normal University(重庆师范大学计算机与信息科学学院) SymtrustAI.com Eastern Institute of Technology, Ningbo(宁波东方理工大学)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 研究发现,不同大型语言模型在相同提示下预测相同目标词时,常共享交互模式,且高级模型一致性更强,共享交互通常阶数更低、正负抵消更弱。

Comments 20 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01609 2026-06-09 cs.CL 版本更新 89%

Swift-SVD: Theoretical Optimality Meets Practical Efficiency in Low-Rank LLM Compression

Swift-SVD:在低秩LLM压缩中理论最优与实用效率的结合

Ruoling Qi, Yirui Liu, Xuaner Wu, Xiangyu Wang, Ming Li, Chen Chen, Jian Chen, Yin Chen, Qizhen Weng

机构 * Nanyang Technological University(南洋理工大学)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出Swift-SVD框架,通过激活感知和闭式压缩方法,在保证理论最优的同时提升实用效率和数值稳定性,实验显示其在压缩精度和端到端压缩时间上有显著优势。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28115 2026-05-28 cs.AI 89%

CIVIC: End-to-End Sequence Compactness for Efficient Vision-Language Models

CIVIC: 面向高效视觉语言模型的端到端序列紧凑性

Fengze Yang, Bo Yu, Xuewen Luo, Cathy Liu, Chenxi Liu

机构 * Department of Civil & Environmental Engineering(土木与环境工程系) University of Utah(犹他大学)

专题命中 效率与部署 :LLM(summary_cn,abstract);language model(title,abstract);分类 cs.AI

AI总结 提出CIVIC框架,通过路径一致的紧凑视觉推理,在视觉编码器、投影层、LLM预填充和KV缓存中保持紧凑序列表示,减少非连续内存访问和局部合并开销,在Qwen3-VL架构上实现KV缓存内存降至约三分之一并降低端到端推理延迟,同时通过文本对齐KL蒸馏和自适应空间保留下限保持精度。

Comments 11 pages, 6 figures, 2 tables, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27763 2026-05-28 cs.LG 89%

A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving

LLM服务中基于批处理条件的拒绝鲁棒性配对测试协议

Sahil Kadadekar

机构 * Independent Researcher(独立研究者)

专题命中 效率与部署 :LLM(title,title_cn);language model(abstract);分类 cs.LG

AI总结 提出配对测试协议,通过四项研究验证批处理条件对LLM安全标签的影响,发现批处理导致的安全标签翻转率低但存在,建议精确堆栈验证。

Comments 12 pages. Accepted to the ICML 2026 Workshop on Hypothesis Testing

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18692 2026-05-28 cs.AI math.OC 89%

Democratizing Large-Scale Re-Optimization with LLM-Guided Model Patches

利用LLM引导的模型补丁实现大规模重新优化的大众化

Tinghan Ye, Arnaud Deza, Ved Mohan, El Mehdi Er Raqabi, Pascal Van Hentenryck

机构 * H. Milton Stewart School of Industrial and Systems Engineering, Georgia Institute of Technology(赫尔曼·米利特·斯图尔特工业与系统工程学院,佐治亚理工学院) Department of Operations and Decision Systems, Université Laval(运营与决策系统系,拉瓦尔大学)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出一个基于大语言模型的代理重新优化框架,通过自然语言交互和优化工具箱,使非专家用户能够动态更新和重新优化部署的优化模型,并在两个大规模实际案例中验证了其有效性和可扩展性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09557 2026-05-21 cs.DC cs.LG 89%

Understanding and Improving Communication Performance in Multi-node LLM Inference

理解并改进多节点LLM推理中的通信性能

Prajwal Singhania, Siddharth Singh, Lannie Dalton Hough, Akarsh Srivastava, Harshitha Menon, Charles Fredrick Jekel, Abhinav Bhatele

机构 * Department of Computer Science, University of Maryland(马里兰大学计算机科学系) Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本研究探讨了多节点分布式推理中通信性能的优化,通过分析不同模型并行方案的强标度行为,提出了一种基于递归倍增的分层all-reduce算法NVRAR,显著降低了推理延迟。

Comments 17 Figures, To Appear in Proceedings of ACM Conference on AI and Agentic Systems 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01752 2026-05-11 cs.CL cs.CR 89%

WorldCup Sampling for Multi-bit LLM Watermarking

多比特LLM水印的WorldCup采样

Yidan Wang, Yubing Ren, Yanan Cao, Li Guo

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络与信息安全学院)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出WorldCup框架,通过结构化通信通道建模和分层竞争机制实现多比特水印嵌入,提升文本质量和解码鲁棒性,实验表明其在容量、检测性、鲁棒性等方面优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27296 2026-05-01 cs.SE cs.CL 89%

To Diff or Not to Diff? Structure-Aware and Adaptive Output Formats for Efficient LLM-based Code Editing

要进行差异生成还是不进行差异生成?面向高效LLM代码编辑的结构感知和自适应输出格式

Wei Cheng, Yongchang Cao, Chen Shen, Binhua Li, Jue Chen, Yongbin Li, Wei Hu

机构 * State Key Laboratory for Novel Software Technology, Nanjing University, China(新型软件技术国家重点实验室,南京大学,中国) Tongyi Lab, Alibaba Group, China(通义实验室,阿里巴巴集团,中国)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出BlockDiff和FuncDiff两种结构感知差异格式及AdaEdit策略,通过减少生成复杂度提升长代码编辑效率,准确度与全代码生成相当,降低30%以上延迟和成本。

Comments Accepted in the Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27032 2026-05-01 cs.SE cs.LG 89%

LLM-Guided Runtime Parameter Optimization for Energy-Efficient Model Inference

基于大语言模型的运行时参数优化以实现高效模型推理

Katelyn Crumpacker, Dimitrios Nikolopoulos

机构 * Virginia Polytechnic and State University(弗吉尼亚理工大学)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文提出一种结合人类反馈和大语言模型的运行时参数优化方法,以提高模型推理的能效,通过迭代优化快速找到高效参数配置。

Comments 8 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14590 2026-04-23 cs.LG 89%

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design

MixLLM: LLM量化与输出特征间全局混合精度的高效系统设计

Zhen Zheng, Xiaonan Song, Chuanjie Liu

机构 * Microsoft(微软)

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.LG

AI总结 MixLLM通过全局混合精度量化优化输出特征,提升LLM压缩精度与系统效率,实验显示在增加10%比特数下, perplexity和MMLU-Pro损失显著降低。

Comments Accepted at MLSys 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13806 2026-04-16 cs.LG 89%

Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate

鲁棒超低比特后训练量化 via 稳定对角曲率估计

Jaemin Kim, Sungkyun Kim, Junyeol Lee, Jiwon Seo

机构 * Seoul National University(首尔国立大学) Hanyang University(翰阳大学)

专题命中 效率与部署 :post-training(title,abstract);LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 本文提出DASH-Q框架,通过使用对角Hessian近似和迭代加权最小二乘法,在超低比特条件下提升零样本准确率,实现更稳定的量化效果。

Comments EUROMLSYS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13556 2026-04-16 cs.CL 89%

YOCO++: Enhancing YOCO with KV Residual Connections for Efficient LLM Inference

YOCO++:通过KV残差连接增强YOCO以实现高效的LLM推理

You Wu, Ziheng Chen, Yizhen Zhang, Haoyi Wu, Chengting Yu, Yuchi Xu, Wenbo Su, Bo Zheng, Kewei Tu

机构 * Alibaba Group(阿里巴巴集团) ShanghaiTech University(上海科技大学) Zhejiang University(浙江大学)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出YOCO++,通过在底层层与底层之间加入加权残差连接,提升YOCO的性能,实现50%的KV缓存压缩率下的最佳表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08003 2026-04-10 eess.AS cs.CL cs.SD 89%

Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs

重新思考基于大语言模型的自动语音识别中的熵分配:理解语音编码器与大语言模型之间的动态关系

Yuan Xie, Jiaqi Song, Guang Qiu, Xianliang Wang, Ming Lei, Jie Gao, Jie Wu

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);pretraining(abstract)

AI总结 本文从熵分配角度重新审视基于大语言模型的ASR,提出三个指标刻画训练范式如何在语音编码器和LLM之间分配熵减少,通过多阶段训练策略优化参数效率和抗幻觉能力,实验显示在2.3B参数下达到与最新模型相当的性能并有效缓解幻觉。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12933 2026-03-16 cs.AI 89%

Efficient and Interpretable Multi-Agent LLM Routing via Ant Colony Optimization

通过蚁群优化实现高效且可解释的多智能体大语言模型路由

Xudong Wang, Chaoning Zhang, Jiaquan Zhang, Chenghao Li, Qigan Sun, Sung-Ho Bae, Peng Wang, Ning Xie, Jie Zou, Yang Yang, Hengtao Shen

机构 * School of Computing, Kyung Hee University(Kyung Hee 大学计算机学院) School of Information and Software Engineering, University of Electronic Science and Technology of China(中国电子科技大学信息与软件工程学院) School of Computer Science and Engineering, University of Electronic Science and Technology of China(中国电子科技大学计算机科学与工程学院) School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);small language model(abstract)

AI总结 本文提出AMRO-S框架,通过意图推理、任务特定信息素专家和质量门控异步更新机制,提升多智能体系统路由效率与可解释性,实验显示其在质量-成本权衡上优于现有方法。

Comments 11 pages, 3 figures, submitted to IEEE Transactions on Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22132 2026-01-30 cs.LG 89%

Pay for Hints, Not Answers: LLM Shepherding for Cost-Efficient Inference

为提示付费,而非答案:用于高效推断的LLM牧师

Ziming Dong, Hardik Sharma, Evan O'Toole, Jaya Prakash Champati, Kui Wu

机构 * Department of Computer Science, University of Victoria, Victoria, Canada(维多利亚大学计算机科学系) Department Of Information Technology, Manipal University, Manipal, India(马那尔大学信息科技系)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);small language model(abstract)

AI总结 LLM牧师通过请求LLM的短提示来提高SLM的准确性,显著降低推断成本,同时保持准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22788 2025-12-01 cs.CR cs.CL 89%

PRISM: Privacy-Aware Routing for Adaptive Cloud-Edge LLM Inference via Semantic Sketch Collaboration

PRISM:通过语义草图协作实现隐私感知的自适应云-边LLM推理

Junfei Zhan, Haoxun Shen, Zheng Lin, Tengjiao He

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);small language model(abstract)

AI总结 PRISM通过语义草图协作实现隐私感知的自适应云-边LLM推理,动态平衡隐私与推理质量,降低能耗和延迟,提升输出质量。

Comments Accepted to AAAI 2026. This is the arXiv preprint version

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17257 2025-10-29 cs.LG q-bio.GN 89%

JanusDNA: A Powerful Bi-directional Hybrid DNA Foundation Model

Qihao Duan, Bingding Huang, Zhenqiao Song, Irina Lehmann, Lei Gu, Roland Eils, Benjamin Wild

机构 * Berlin Institute of Health, Charité – Universitätsmedizin Berlin(柏林健康研究所,柏林夏里特大学医学中心) College of Big Data and Internet, Shenzhen Technology University(大数据与互联网学院,深圳技术大学) Language Technologies Institute, Carnegie Mellon University(语言技术研究所,卡内基梅隆大学) Epigenetics Laboratory, Max Planck Institute for Heart and Lung Research(表观遗传学实验室,马克斯·普朗克心脏病和肺部研究所以) Department of Mathematics and Computer Science, Freie Universität Berlin(数学与计算机科学系,柏林自由大学) Health Data Science Unit, Heidelberg University Hospital and BioQuant(健康数据科学单元,海德堡大学医院和BioQuant) Intelligent Medicine Institute, Fudan University(智能医学研究所,复旦大学)

专题命中 效率与部署 :foundation model(title,abstract);LLM(abstract);large language model(abstract);language model(abstract)

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02133 2025-09-03 cs.CL 89%

AMBEDKAR-A Multi-level Bias Elimination through a Decoding Approach with Knowledge Augmentation for Robust Constitutional Alignment of Language Models

Snehasis Mukhopadhyay, Aryan Kasat, Shivam Dubey, Rahul Karthikeyan, Dhruv Sood, Vinija Jain, Aman Chadha, Amitava Das

机构 * Indian Institute of Information Technology, Kalyani(印度信息技术学院,卡里尼) BITS Pilani Goa(比尔·斯图尔特学院,果阿) IIT Madras(马德拉斯理工学院) DTU(达丁理工大学) Artificial Intelligence Institute, University of South Carolina(南卡罗来纳大学人工智能研究所) Meta AI Amazon GenAI(亚马逊生成人工智能)

专题命中 效率与部署 :language model(title,abstract);LLM(abstract);large language model(abstract);small language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06402 2025-08-08 cs.RO cs.AI cs.HC 89%

Camera Control at the Edge with Language Models for Scene Understanding

Alexiy Buynitsky, Sina Ehsani, Bhanu Pallakonda, Pragyana Mishra

机构 * Purdue University(普渡大学)

专题命中 效率与部署 :language model(title,abstract);LLM(abstract);large language model(abstract);SFT(abstract)

Comments 7 pages, 6 figures. This work was presented and published at the 11th IEEE International Conference on Control, Automation and Robotics (ICCAR) in 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16056 2025-04-23 cs.CL 89%

Honey, I Shrunk the Language Model: Impact of Knowledge Distillation Methods on Performance and Explainability

Daniel Hendriks, Philipp Spitzer, Niklas Kühl, Gerhard Satzger

机构 * Karlsruhe Institute of Technology (KIT)(卡尔斯鲁厄理工学院) Institute for Information Systems (WIN)(信息系统研究所) University of Bayreuth(拜罗伊特大学)

专题命中 效率与部署 :language model(title,abstract);LLM(abstract);large language model(abstract);small language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12167 2025-03-20 cs.CL 89%

PLM: Efficient Peripheral Language Models Hardware-Co-Designed for Ubiquitous Computing

Cheng Deng, Luoyang Sun, Jiwen Jiang, Yongcheng Zeng, Xinjian Wu, Wenxin Zhao, Qingfa Xiao, Jiachuan Wang, Haoyang Li, Lei Chen, Lionel M. Ni, Haifeng Zhang, Jun Wang

专题命中 效率与部署 :language model(title,abstract);large language model(abstract);small language model(abstract);SFT(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11499 2024-12-17 cs.AI cs.RO 89%

Embodied CoT Distillation From LLM To Off-the-shelf Agents

Wonje Choi, Woo Kyung Kim, Minjong Yoo, Honguk Woo

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);small language model(abstract)

Comments Accepted at ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09758 2024-10-01 cs.CL 89%

OrchestraLLM: Efficient Orchestration of Language Models for Dialogue State Tracking

Chia-Hsuan Lee, Hao Cheng, Mari Ostendorf

专题命中 效率与部署 :language model(title,abstract);LLM(abstract);large language model(abstract);small language model(abstract)

Comments updated version (NAACL camera ready)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06668 2024-03-19 cs.LG cs.CV 89%

Large Language Models and Foundation Models in Smart Agriculture: Basics, Opportunities, and Challenges

Jiajia Li, Mingle Xu, Lirong Xiang, Dong Chen, Weichao Zhuang, Xunyuan Yin, Zhaojian Li

专题命中 效率与部署 :large language model(title);language model(title);foundation model(title);分类 cs.LG

Comments 18 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06801 2024-02-20 cs.AI 89%

Graph-of-Thought: Utilizing Large Language Models to Solve Complex and Dynamic Business Problems

Ye Li

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

Comments Keywords: Graph-of-Thought (GoT), Workflow Automation, Large Language Models (LLMs), Task Execution, Data-Driven Decision Making, Complexity Management

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00819 2023-10-03 cs.CL 89%

Parameter-Efficient Tuning Helps Language Model Alignment

Tianci Xue, Ziqi Wang, Heng Ji

专题命中 效率与部署 :language model(title,abstract);large language model(abstract);RLHF(abstract);preference optimization(abstract)

Comments 21 pages, 11 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07093 2024-07-10 cs.CL cs.AI cs.LG 89%

FBI-LLM: Scaling Up Fully Binarized LLMs from Scratch via Autoregressive Distillation

Liqun Ma, Mingjie Sun, Zhiqiang Shen

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);pretraining(abstract)

Comments Github at https://github.com/LiqunMa/FBI-LLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07725 2026-08-19 cs.CL cs.AI 版本更新 88%

SOD: Step-wise On-policy Distillation for Small Language Model Agents

SOD:分步式在线蒸馏用于小型语言模型代理

Qiyong Zhong, Mao Zheng, Mingyang Song, Xin Lin, Jie Sun, Houcheng Jiang, Xiang Wang, Junfeng Fang

机构 * Zhejiang University(浙江大学) Large Language Model Department, Tencent(腾讯大语言模型部门) University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学)

专题命中 效率与部署 :language model(title,abstract);small language model(title,abstract);分类 cs.CL、cs.AI

AI总结 针对小型语言模型中工具集成推理的稳定性问题,提出SOD分步式在线蒸馏框架,通过动态调整蒸馏强度缓解教师信号误导,提升推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14644 2026-08-18 cs.LG cs.CL 新提交 88%

DUET: Dual-Teacher On-Policy Distillation via Same-Weight Disagreement for Prohibition Compliance

DUET:基于同权重分歧的双教师在线策略蒸馏用于禁止合规性

Zihan Li, Feifei Li, Wenhui Que

专题命中 效率与部署 :LLM(summary_cn,abstract);SFT(abstract,abstract_cn);post-training(abstract);分类 cs.CL、cs.LG

AI总结 本研究针对LLM部署中的动态禁止规则合规问题,提出DUET双教师在线策略蒸馏方法,构建工业基准,在Qwen模型上实现高合规性与效用保留,性能优于基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14604 2026-08-18 cs.CL cs.AI 新提交 88%

Wiola 13M, a Gated Spiral Attention Architecture for Parameter Efficient Small Language Models

Wiola 13M:一种用于参数高效小型语言模型的门控螺旋注意力架构

Aryuemaan Kumar Chowdhury, Praveen Oosa, Vineesha Reddy

专题命中 效率与部署 :language model(title,abstract);small language model(title,abstract);分类 cs.CL、cs.AI

AI总结 针对10M-100M参数小型语言模型未适配小规模Transformer的问题,提出含三个创新组件的Wiola 13M模型,验证其等价性并发布开源实现,提升长程区分与梯度流性能。

Comments 6

详情

展开后加载摘要…

URL PDF HTML 收藏