arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-08-14 至 2025-08-14 共收录 25 信号源:cs.CL, cs.AI, cs.LG

1. 效率与部署 25 篇

2508.09471 2025-08-14 cs.LG 92%

EGGS-PTP: An Expander-Graph Guided Structured Post-training Pruning Method for Large Language Models

Omar Bazarbachi, Zijun Sun, Yanning Shen

机构 * Department of Electrical Engineering and Computer Science University of California, Irvine(电气工程与计算机科学系加州大学伊文斯顿分校)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);post-training(title,abstract);foundation model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09834 2025-08-14 cs.CL cs.AI cs.CV 91%

Speed Always Wins: A Survey on Efficient Architectures for Large Language Models

Weigao Sun, Jiaxi Hu, Yucheng Zhou, Jusen Du, Disen Lan, Kexin Wang, Tong Zhu, Xiaoye Qu, Yu Zhang, Xiaoyu Mo, Daizong Liu, Yuxuan Liang, Wenliang Chen, Guoqi Li, Yu Cheng

机构 * Shanghai AI Laboratory(上海人工智能实验室) HKUST (GZ)(香港科技大学) University of Macau(澳门大学) Institute of Automation Chinese Academy of Sciences(中国科学院自动化研究所) Soochow University(苏州大学) KTH Royal Institute of Technology(皇家理工学院) Peking University(北京大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract);foundation model(abstract)

Comments Survey, 82 pages, GitHub: https://github.com/weigao266/Awesome-Efficient-Arch

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19090 2025-08-14 cs.LG 88%

Pivoting Factorization: A Compact Meta Low-Rank Representation of Sparsity for Efficient Inference in Large Language Models

Jialin Zhao, Yingtao Zhang, Carlo Vittorio Cannistraci

机构 * Center for Complex Network Intelligence (CCNI), Tsinghua Laboratory of Brain and Intelligence (THBI), Department of Psychological and Cognitive Sciences(复杂网络智能中心(CCNI)、清华脑智能实验室(THBI)、心理与认知科学系) Department of Computer Science(计算机科学系) Department of Biomedical Engineering, Tsinghua University, China(生物医学工程系,清华大学,中国)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);分类 cs.LG

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14910 2025-08-14 cs.CL cs.AI 86%

EvoP: Robust LLM Inference via Evolutionary Pruning

Shangyu Wu, Hongchao Du, Ying Xiong, Shuai Chen, Tei-Wei Kuo, Nan Guan, Chun Jason Xue

机构 * City University of Hong Kong(香港城市大学) Mohamed bin Zayed University of Artificial Intelligence(马尔代夫穆罕默德·本·扎耶德人工智能大学) Baidu(百度) National Taiwan University(国立台湾大学)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09173 2025-08-14 cs.NI 85%

Camel: Energy-Aware LLM Inference on Resource-Constrained Devices

Hao Xu, Long Peng, Shezheng Song, Xiaodong Liu, Ma Jun, Shasha Li, Jie Yu, Xiaoguang Mao

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09149 2025-08-14 cs.NI cs.DC 85%

Semantic-Aware LLM Orchestration for Proactive Resource Management in Predictive Digital Twin Vehicular Networks

Seyed Hossein Ahmadpanah

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09148 2025-08-14 cs.LG cs.AI 82%

Motif 2.6B Technical Report

Junghwan Lim, Sungmin Lee, Dongseok Kim, Eunhwan Park, Hyunbyung Park, Junhyeok Lee, Wai Ting Cheung, Dahye Choi, Jaeheui Her, Jaeyeon Huh, Hanbin Jung, Changjin Kang, Beomgyu Kim, Jihwan Kim, Minjae Kim, Taehwan Kim, Youngrok Kim, Haesol Lee, Jeesoo Lee, Kungyu Lee, Dongpin Oh, Yeongjae Park, Bokki Ryu, Daewon Suh, Dongjoo Weon

机构 * Motif Technologies

专题命中 效率与部署 :LLM(abstract);large language model(abstract);language model(abstract);foundation model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09240 2025-08-14 cs.NI cs.AI cs.CL 79%

NEFMind: Parameter-Efficient Fine-Tuning of Open-Source LLMs for Telecom APIs Automation

Zainab Khan, Ahmed Hussain, Mukesh Thakur, Arto Hellas, Panos Papadimitratos

机构 * Alto University, School of Science, Espoo, Finland(阿尔托大学科学学院,芬兰) Networked Systems Security (NSS) Group -- KTH Royal Institute of Technology, Stockholm, Sweden(网络系统安全(NSS)小组——皇家理工学院,瑞典) Ericsson, Finland(爱立信,芬兰)

专题命中 效率与部署 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09199 2025-08-14 cs.CV cs.AI cs.CL 79%

$Δ$-AttnMask: Attention-Guided Masked Hidden States for Efficient Data Selection and Augmentation

Jucheng Hu, Suorong Yang, Dongzhan Zhou

专题命中 效率与部署 :large language model(abstract);language model(abstract);post-training(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06323 2025-08-14 cs.LG cs.AI 79%

Mosaic: Composite Projection Pruning for Resource-efficient LLMs

Bailey J. Eccles, Leon Wong, Blesson Varghese

机构 * organization= School of Computer Science, University of St Andrews , addressline= Jack Cole Building , city= St Andrews , postcode= KY16 9SX , state= Fife , country= Scotland, United Kingdom organization= Autonomous Networking Research \& Innovation Department, Rakuten Mobile, Inc. , addressline= Rakuten Crimson House, 1-14-1 Tamagawa, Setagaya-ku , city= Tokyo , postcode= 158-0094 , state= Tokyo , country= Japan

专题命中 效率与部署 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09470 2025-08-14 cs.CV 78%

CitySeg: A 3D Open Vocabulary Semantic Segmentation Foundation Model in City-scale Scenarios

Jialei Xu, Zizhuang Wei, Weikang You, Linyun Li, Weijian Sun

机构 * Huawei Technologies Co., Ltd(华为技术有限公司)

专题命中 效率与部署 :foundation model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09500 2025-08-14 cs.CV 78%

Advancing Reliable Test-Time Adaptation of Vision-Language Models under Visual Variations

Yiwen Liang, Hui Chen, Yizhe Xiong, Zihan Zhou, Mengyao Lyu, Zijia Lin, Shuaicheng Niu, Sicheng Zhao, Jungong Han, Guiguang Ding

机构 * Tsinghua University(清华大学) Nanyang Technological University(南洋理工大学)

专题命中 效率与部署 :language model(title,abstract)

Comments Accepted at the 33rd ACM International Conference on Multimedia(ACM MM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09755 2025-08-14 cs.CL 77%

Transforming Questions and Documents for Semantically Aligned Retrieval-Augmented Generation

Seokgi Lee

机构 * Seokgi Lee

专题命中 效率与部署 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09194 2025-08-14 cs.LG cs.AI 73%

Meta-Learning for Speeding Up Large Model Inference in Decentralized Environments

Yipeng Du, Zihao Wang, Ahmad Farhan, Claudio Angione, Harry Yang, Fielding Johnston, James P. Buban, Patrick Colangelo, Yue Zhao, Yuzhe Yang

机构 * Nesa Research(Nesa研究机构)

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

Comments COLM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09192 2025-08-14 cs.LG cs.AI 73%

Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing

Xu Wang, Chenkai Xu, Yijie Jin, Jiachun Jin, Hao Zhang, Zhijie Deng

机构 * Shanghai Jiao Tong University(上海交通大学) University of California San Diego(加州大学圣地亚哥分校) Shanghai University(上海大学)

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23077 2025-08-14 cs.CL 70%

Efficient Inference for Large Reasoning Models: A Survey

Yue Liu, Jiaying Wu, Yufei He, Ruihan Gong, Jun Xia, Liang Li, Hongcheng Gao, Hongyu Chen, Baolong Bi, Jiaheng Zhang, Zhiqi Huang, Bryan Hooi, Stan Z. Li, Keqin Li

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09208 2025-08-14 cs.NI cs.AI 70%

CoMoE: Collaborative Optimization of Expert Aggregation and Offloading for MoE-based LLMs at Edge

Muqing Li, Ning Li, Xin Yuan, Wenchao Xu, Quan Chen, Song Guo, Haijun Zhang

机构 * Harbin Institute of Technology(哈尔滨工业大学) Hong Kong University of Science and Technology(香港科技大学) Guangdong University of Technology(广东工业大学) University of Science and Technology Beijing(北京科技大学)

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07315 2025-08-14 eess.AS cs.AI cs.CL cs.LG cs.SD 67%

FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities

Lilit Grigoryan, Vladimir Bataev, Nikolay Karpov, Andrei Andrusenko, Vitaly Lavrukhin, Boris Ginsburg

机构 * NVIDIA

专题命中 效率与部署 :language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to Automatic Speech Recognition and Understanding Workshop (ASRU) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20237 2025-08-14 cs.CL cs.SD eess.AS 57%

Efficient Speech Translation through Model Compression and Knowledge Distillation

Yasmin Moslem

机构 * ADAPT Centre School of Computer Science and Statistics(ADAPT中心计算机科学与统计学学院)

专题命中 效率与部署 :language model(abstract);分类 cs.CL

Comments IWSLT 2025

Journal ref Proceedings of the 22nd International Conference on Spoken Language Translation (IWSLT 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09561 2025-08-14 cs.LG 57%

Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges

Changyuan Zhao, Guangyuan Liu, Ruichen Zhang, Yinqiu Liu, Jiacheng Wang, Jiawen Kang, Dusit Niyato, Zan Li, Xuemin, Shen, Zhu Han, Sumei Sun, Chau Yuen, Dong In Kim

机构 * College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院) School of Automation, Guangdong University of Technology(广东工业大学自动化学院) State Key Laboratory of Integrated Services Networks, Xidian University(西安电子科技大学集成服务网络国家重点实验室) Department of Electrical and Computer Engineering, University of Waterloo(滑铁卢大学电气与计算机工程系) Department of Computer Science and Engineering, Kyung Hee University(韩国庆熙大学计算机科学与工程系) Institute for Infocomm Research, Agency for Science, Technology and Research(科技研究局信息通信研究所) Department of Electrical and Computer Engineering, Sungkyunkwan University(庆熙大学电气与计算机工程系)

专题命中 效率与部署 :foundation model(abstract);分类 cs.LG

Comments 21 pages. 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09500 2025-08-14 cs.LG cs.AR 57%

MiCo: End-to-End Mixed Precision Neural Network Co-Exploration Framework for Edge AI

Zijun Jiang, Yangdi Lyu

机构 * Microelectronics Thrust, The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)微电子研究组)

专题命中 效率与部署 :post-training(abstract);分类 cs.LG

Comments 9 pages, 6 figures, accepted by ICCAD'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09499 2025-08-14 cs.CV cs.CG cs.LG 57%

CWFBind: Geometry-Awareness for Fast and Accurate Protein-Ligand Docking

Liyan Jia, Chuan-Xian Ren, Hong Yan

机构 * School of Mathematics, Sun Yat-Sen University(中山大学数学学院) Department of Electrical Engineering, City University of Hong Kong(香港城市大学电子工程系)

专题命中 效率与部署 :language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09229 2025-08-14 cs.NI cs.AI cs.DC 57%

Cluster Topology-Driven Placement of Experts Reduces Network Traffic in MoE Inference

Danil Sivtsov, Aleksandr Katrutsa, Ivan Oseledets

机构 * AIRI Skoltech(斯克里普钦斯基理工学院)

专题命中 效率与部署 :LLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09197 2025-08-14 cs.NI cs.AI 57%

MX-AI: Agentic Observability and Control Platform for Open and AI-RAN

Ilias Chatzistefanidis, Andrea Leone, Ali Yaghoubian, Mikel Irazabal, Sehad Nassim, Lina Bariah, Merouane Debbah, Navid Nikaein

机构 * Communications Department(通讯部门) EURECOM AI Department(AI部门) BubbleRAN RIC Department(RIC部门) Aalto University(阿莱大学) EECS Department(EECS部门) Khalifa University(卡利法大学)

专题命中 效率与部署 :LLM(abstract);分类 cs.AI

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09366 2025-08-14 cs.SE 50%

Plug it and Play on Logs: A Configuration-Free Statistic-Based Log Parser

Qiaolin Qin, Xingfang Wu, Heng Li, Ettore Merlo

专题命中 效率与部署 :language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏