Comments5 pages, 1 figures. Accepted at the 2nd International Workshop on Low Carbon Computing (LOCO 2026), Lancaster University, United Kingdom, 10-11 September 2026. Part of the LOCO 2026 proceedings, arXiv: LOCO2026/P14
Abstract representational geometry supports inference in large language models
抽象表征几何支持大型语言模型中的推理
Yunan Zeng, Yuwang Wang
机构
*
College of Future Information Technology, Fudan University(复旦大学未来信息科技学院)
;
Beijing National Research Center for Information Science and Technology, Tsinghua University(北京信息科学与技术国家研究中心,清华大学)
专题命中
效率与部署
:large language model(title,abstract);language model(title,abstract);LLM(summary_cn,abstract_cn);分类 cs.AI
TailSieve: Partial-Rollout-Guided Tail Routing for LLM Rollouts
TailSieve:面向大语言模型(LLM)rollout的部分rollout引导式长尾路由
Tianqi Xu, Lu Lv, Haoyang Huang, Wenjie Huang, Zhanming Shen, Yuhao Shen, Baolin Zhang, Xinyi Hu, Shuang Ge, Jun Dai, Tianyu Liu, Suorong Yang, Zhikai Li, Ye Bai, Jun Zhang, Lei Chen, Yue Li, Mingchen Wan
机构
*
Qwen Business Unit of Alibaba(阿里巴巴通义千问业务部)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Zhejiang University(浙江大学)
;
University of Science and Technology of China(中国科学技术大学)
;
National University of Singapore(新加坡国立大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization
迈向高效大语言模型服务:关于系统感知键值缓存优化的综述
Jiantong Jiang, Peiyu Yang, Rui Zhang, Feng Liu
机构
*
School of Computing and Information Systems, The University of Melbourne(墨尔本大学计算与信息系统学院)
;
School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院)
专题命中
效率与部署
:large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.LG