POP: Online Structural Pruning Enables Efficient Inference of Large Foundation Models
POP:在线结构剪枝实现大基础模型的高效推理
Yi Chen, Wonjin Shin, Shuhong Liu, Tho Mai, Jeongmo Lee, Chuanbo Hua, Kun Wang, Jun Liu, Joo-Young Kim
机构
*
Korea Advanced Institute of Science and Technology(韩国科学技术院)
;
University of Tokyo, Tokyo, Japan(东京大学)
;
Tokyo Institute of Technology, Tokyo, Japan(东京技术大学)
专题命中
效率与部署
:foundation model(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI
Bridging 6G IoT and AI: LLM-Based Efficient Approach for Physical Layer's Optimization Tasks
连接6G物联网与AI:基于大语言模型的高效方法用于物理层优化任务
Ahsan Mehmood, Naveed Ul Hassan, Ghassan M. Kraidy
机构
*
Department of Electronic Systems, Norwegian University of Science and Technology(电子系统系,挪威科学技术大学)
;
Department of Electrical Engineering, Lahore University of Management sciences (LUMS)(电气工程系,拉合尔管理科学大学(LUMS))
专题命中
效率与部署
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI
PackInfer: Compute- and I/O-Efficient Attention for Batched LLM Inference
PackInfer: 用于批量LLM推理的计算和I/O高效注意力
Rui Ning, Wei Zhang, Fan Lai
机构
*
Nanjing University, Nanjing, China(南京大学)
;
Siebel Center for Computer Science, University of Illinois Urbana-Champaign, Urbana, IL, USA(伊利诺伊大学厄巴纳-香槟分校计算机科学中心)
专题命中
效率与部署
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.LG
机构
*
School of Internet of Things Engineering, Jiangnan University(江南大学物联网工程学院)
;
School of Information Engineering, Jiangxi Provincial Key Laboratory of Advanced Signal Processing and Intelligent Communications, Nanchang University(江西省级先进信号处理与智能通信重点实验室,南昌大学信息工程学院)
;
Department of Electronic Engineering, State Key Laboratory of Space Network and Communications, and the Beijing National Research Center for Information Science and Technology, Tsinghua University(电子工程系,空间网络与通信国家重点实验室,信息科学与技术国家研究中心,清华大学)
;
Department of Computer Science, Brunel University(计算机科学系,布鲁内尔大学)
;
State Key Laboratory of ISN and the School of Telecommunications Engineering, Xidian University(信息与通信国家重点实验室,西安电子科技大学电信工程学院)
专题命中
效率与部署
:LLM(title);large language model(abstract);language model(abstract);prompting(abstract)
ELLMPEG: An Edge-based Agentic LLM Video Processing Tool
ELLMPEG:一种基于边缘的代理LLM视频处理工具
Zoha Azimi, Reza Farahani, Radu Prodan, Christian Timmerer
机构
*
Christian Doppler Laboratory ATHENA, Department of Information Technology (ITEC) University of Klagenfurt(亚琛实验室ATHENA,信息科技系,克雷格弗尔特大学)
;
Department of Information Technology (ITEC) University of Klagenfurt(信息科技系,克雷格弗尔特大学)
;
Department of Computer Science University of Innsbruck(计算机科学系,因斯布鲁克大学)
专题命中
效率与部署
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.LG
Stackelberg Self-Annotation: A Robust Approach to Data-Efficient LLM Alignment
Stackelberg 自注释:一种鲁棒的数据高效 LLM 对齐方法
Xu Chu, Zhixin Zhang, Tianyu Jia, Yujie Jin
机构
*
Key Laboratory of High Confidence Software Technologies, Ministry of Education(高可信软件技术重点实验室,教育部)
;
Center on Frontiers of Computing Studies, Peking University(计算前沿研究中心,北京大学)
;
School of Computer Science, Peking University(计算机学院,北京大学)
专题命中
效率与部署
:LLM(title);large language model(abstract);language model(abstract);preference optimization(abstract)
机构
*
Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系)
;
Division of Integrative Systems and Design, The Hong Kong University of Science and Technology(香港理工大学系统与设计学院)
;
School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院)
;
Department of Electrical and Computer Engineering, University of Waterloo(滑铁卢大学电气与计算机工程系)
专题命中
效率与部署
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI
AI总结
HALO通过语义感知预测和负载平衡调度,提升失真边缘网络中LLM推理的效率与性能。
CommentsAccepted by IEEE International Conference on Computer Communications (INFOCOM) 2026