Comments8 pages, 4 tables, 1 algorithm. v2: corrects MLA compression to 57x (576 dims per DeepSeek); fixes a factor-of-two in GQA sizing rows and all dependent transfer numbers (Llama-70B 4K = 1.3 GB); pipelining band recomputed (76 to 100 percent); PCIe on Gen5; NIXL positioning added; NVLink 5 and 2026 model notes (bandwidth- and bytes-parametric); wording fixes
Tianjin Huang, Zhangyang Wang, Haotian Hu, Zhenyu Zhang, Gaojie Jin, Xiang Li, Li Shen, Jiaxing Shang, Tianlong Chen, Ke Li, Lu Liu, Qingsong Wen, Shiwei Liu
机构
*
Department of Computer Science, University of Exeter(埃克塞特大学计算机科学系)
;
Department of Mathematics and Computer Science, Eindhoven University of Technology(埃因霍温理工大学数学与计算机科学系)
;
School of the Gifted Young, University of Science and Technology of China(中国科学技术大学天才青年学院)
;
Department of Electrical and Computer Engineering, University of Texas at Austin(德克萨斯大学奥斯汀分校电气与计算机工程系)
;
Department of Computer Science, University of Reading(阅读大学计算机科学系)
;
School of Cyber Science and Technology, Sun Yat-sen University(中山大学网络科学与技术学院)
;
Department of Computer Science, The University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校计算机科学系)
;
ELLIS Institute Tubingen(图宾根ELLIS研究所)
;
Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所)
;
Tübingen AI Center, Tübingen, Germany(图宾根人工智能中心,德国图宾根)
;
College of Computer Science, Chongqing University(重庆大学计算机学院)
Comments12 pages, 4 figures, 3 tables. v3 corrects a differential length bias in the automated correctness scorer, validated against 1,830 human adjudications; all analyses recomputed. The inverse accuracy-efficiency coupling in v1/v2 does not survive relabelling; see the version note on page 1. Pre-registered. Code and annotations: github.com/synthiumjp/metacognition-audit
GoQuant: Geometric Orthogonal Residual Projection for Multiplier-Free Power-of-Two Transformer Quantization
GoQuant: 用于无乘法器二次幂变压器量化的几何正交残差投影
Maoyang Xiang, Tao Luo, Bo Wang
机构
*
Information Systems Technology and Design(信息系统技术与设计)
;
Singapore University of Technology and Design(新加坡科技设计大学)
;
Institute of High Performance Computing (IHPC)(高性能计算研究所)
;
Agency for Science, Technology and Research (A*STAR)(科技研究局)
专题命中
效率与部署
:large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
机构
*
City University of Hong Kong(香港城市大学)
;
Tsinghua University(清华大学)
;
Shenzhen University of Advanced Technology(深圳理工大学)
;
Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
专题命中
效率与部署
:large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
Toward Efficient Agents: Memory, Tool learning, and Planning
迈向高效智能体:记忆、工具学习与规划
Xiaofang Yang, Lijun Li, Heng Zhou, Tong Zhu, Xiaoye Qu, Yuchen Fan, Qianshan Wei, Rui Ye, Li Kang, Yiran Qin, Daizong Liu, Qi Li, Ning Ding, Siheng Chen, Jing Shao
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Fudan University(复旦大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Shanghai Jiaotong University(上海交通大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
The Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳))
;
Hong Kong Polytechnic University(香港理工大学)
;
Wuhan University(武汉大学)
;
Tsinghua University(清华大学)
专题命中
效率与部署
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
Comments15 pages, 3 figures, 11 tables. Accepted to HPDC '26 (35th International Symposium on High-Performance Parallel and Distributed Computing), July 13-16, 2026, Cleveland, OH, USA
机构
*
University of Southern California(南加州大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Mohamed Bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
专题命中
效率与部署
:large language model(abstract);language model(abstract);分类 cs.CL、cs.LG
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
;
University of Science and Technology of China(中国科学技术大学)
;
Zhejiang Lab(浙江实验室)
;
Peng Cheng Laboratory(鹏城实验室)