Exploring Layer-wise Information Effectiveness for Post-Training Quantization in Small Language Models
探索分层信息有效性以实现小语言模型的后训练量化
He Xiao, Qingyao Yang, Dirui Xie, Wendong Xu, Zunhai Su, Runming yang, Wenyong Zhou, Haobo Liu, Zhengwu Liu, Ngai Wong
机构
*
The University of Hong Kong(香港大学)
;
Huazhong University of Science and Technology(华中科技大学)
;
Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生学院)
专题命中
效率与部署
:language model(title,abstract);small language model(title,abstract);post-training(title,abstract);large language model(abstract)
DySK-Attn: A Framework for Efficient, Real-Time Knowledge Updating in Large Language Models via Dynamic Sparse Knowledge Attention
DySK-Attn:通过动态稀疏知识注意力实现大语言模型高效实时知识更新的框架
Kabir Khan, Priya Sharma, Arjun Mehta, Neha Gupta, Ravi Narayanan
机构
*
Department of Computer Science, San Francisco State University, San Francisco, CA 94132, India(计算机科学系,圣何塞州立大学)
;
Department of Computer Science and Engineering, Indian Institute of Technology Bombay, Mumbai 400076, India(印度班加罗尔理工学院计算机科学与工程系)
;
Department of Computer Science and Engineering, Indian Institute of Technology Delhi, New Delhi 110016, India(印度德里理工学院计算机科学与工程系)
;
Department of Computer Science and Automation, Indian Institute of Science, Bengaluru 560012, India(印度班加罗尔科学研究所计算机科学与自动化系)
;
Machine Learning Lab, International Institute of Information Technology Hyderabad (IIIT-H), Hyderabad 500032, India(国际信息科技研究所海得拉巴分所机器学习实验室)
专题命中
效率与部署
:large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL、cs.AI、cs.LG
Optimizing Resource Allocation for Geographically-Distributed Inference by Large Language Models
通过大语言模型进行地理分布式推断的资源分配优化
Tingyang Sun, Ting He, Bo Ji, Parimal Parag
机构
*
Department of Computer Science and Engineering, Pennsylvania State University, University Park, PA, USA(计算机科学与工程系,宾夕法尼亚州立大学,大学公园,PA,USA)
;
Department of Computer Science, Virginia Tech, Blacksburg, VA, USA(计算机科学系,弗吉尼亚理工大学,布莱克堡,VA,USA)
;
Department of Electrical Communication Engineering, Indian Institute of Science, Bangalore, India(电子通信工程系,印度科学研究院,班加罗尔,印度)
专题命中
效率与部署
:large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI
Computational Economics in Large Language Models: Exploring Model Behavior and Incentive Design under Resource Constraints
在大型语言模型中进行计算经济学:在资源限制下探索模型行为和激励设计
Sandeep Reddy, Kabir Khan, Rohit Patil, Ananya Chakraborty, Faizan A. Khan, Swati Kulkarni, Arjun Verma, Neha Singh
机构
*
Department of Computer Science and Engineering, Jorhat Engineering College(计算机科学与工程系,贾尔哈特工程学院)
;
School of Computer Science, KLE Technological University(计算机科学学院,KLE技术大学)
;
Department of Computer Applications, Bundelkhand University(计算机应用系,邦德尔坎德大学)
;
Department of Computer Science, Sant Gadge Baba Amravati University(计算机科学系,桑塔·加德·巴瓦·阿姆拉瓦蒂大学)
;
San Francisco State University(旧金山州立大学)
专题命中
效率与部署
:large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL
CommentsPreprint; 7 figures, 4 tables, 1 algorithm. Experiments on GLUE (MNLI, STS-B, CoLA) and WikiText-103 with BERT-base; evaluation includes FLOPS, latency, Gini and entropy metrics
Operationalizing a Threat Model for Red-Teaming Large Language Models (LLMs)
为大规模语言模型(LLMs)的红队测试构建威胁模型
Apurv Verma, Satyapriya Krishna, Sebastian Gehrmann, Madhavan Seshadri, Anu Pradhan, Tom Ault, Leslie Barrett, David Rabinowitz, John Doucette, NhatHai Phan
专题命中
效率与部署
:large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL
机构
*
Graduate School of Life Science and Systems Engineering, Kyushu Institute of Technology, Japan(九州工学技术大学生命科学与系统工程研究生院)
;
Research Center for Neuromorphic AI Hardware, Kyushu Institute of Technology, Japan(九州工学技术大学神经形态人工智能硬件研究中心)
专题命中
效率与部署
:language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI
Comments11 pages, 9 figures. Accepted by ACM for presentation at UCC '25 (18th International Conference on Utility and Cloud Computing), December 1-4, 2025, France. Proceedings publication pending
Ander Alvarez, Alessandro Genuardi, Nilotpal Sinha, Antonio Tiene, Mikail Okyay, Bakbergen Ryskulov, David Montero, Samuel Mugel, Román Orús
机构
*
Multiverse Computing
;
Donostia International Physics Center
;
Ikerbasque Foundation for Science
;
Multiverse Computing, Centre for Social Innovation
专题命中
效率与部署
:large language model(abstract);language model(abstract);分类 cs.AI