LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
LLMRouter:用于开发、评估和部署LLM路由器的统一基础设施
Tao Feng, Fangxu Yu, Haozhen Zhang, Zhongjie Dai, Liangqi Yuan, Zijie Lei, Weizhi Zhang, Kunlun Zhu, Haodong Yue, Keyang Xuan, Ge Liu, Jiaxuan You
机构
*
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
University of Maryland, College Park(马里兰大学帕克分校)
;
Nanyang Technological University(南洋理工大学)
;
Purdue University(普渡大学)
;
University of Illinois Chicago(芝加哥伊利诺伊大学)
专题命中
效率与部署
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL
From Tokens to Watt-hours: Analytical Energy Estimation for LLM Inference on Modern GPUs
从Token到瓦时:现代GPU上大语言模型(LLM)推理的分析式能耗估算
Tina Vartziotis, Rodopi Kosteli, Elli Vartziotis, George Dasoulas, Michael Keckeisen, Konstantinos Skianis, Sotirios Kotsopoulos, Francesca Dominici
机构
*
National Technical University of Athens(雅典国家技术大学)
;
Harvard University(哈佛大学)
;
TWT GmbH Science & Innovation(TWT有限责任公司科学与创新部)
;
NIKI Ltd Digital Engineering(尼基数字工程有限公司)
;
University of Ioannina(约阿尼纳大学)
;
National and Kapodistrian University of Athens(雅典国立卡波迪斯特里亚大学)
;
Massachusetts Institute of Technology(麻省理工学院)
专题命中
效率与部署
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.LG
A Tutorial on Autonomous Fault-Tolerant Control Using Knowledge-Grounded LLM Agents
基于知识接地LLM代理的自主容错控制教程
Javal Vyas, Milapji Singh Gill, Artan Markaj, Felix Gehlhoff, Mehmet Mercangöz
机构
*
Autonomous Industrial Systems Laboratory, Imperial College London(帝国理工学院自主工业系统实验室)
;
Institute of Automation Technology, Helmut Schmidt University(赫尔穆特·施密特大学自动化技术研究所)
专题命中
效率与部署
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI
Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice
面向LLM服务的几何感知在线调度:从理论界到系统实践
Li Kong, Qi Qi, Yinyu Ye, Zijie Zhou
机构
*
Gaoling School of Artificial Intelligence(人工智能学院)
;
Renmin University of China(中国人民大学)
;
Department of Management Science and Engineering(管理科学与工程系)
;
Stanford University(斯坦福大学)
;
Department of Industrial Engineering and Decision Analytics(工业工程与决策分析系)
;
HKUST(香港科技大学)
专题命中
效率与部署
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI
AI总结
针对LLM推理中KV缓存动态内存占用问题,提出Smallest Volume First (SVF)算法,通过几何感知调度优化性能,理论证明将竞争比从48降至5,并在vLLM中实现即插即用,显著降低平均和尾部延迟。
Sim2Schedule: A Simulator-Guided LLM Framework for Autonomous Open-Pit Mine Scheduling
Sim2Schedule: 一种模拟器引导的LLM框架用于自主露天矿调度
Mustavi Ibne Masum, Thiago Eustaquio Alves de Oliveira, Mahzabeen Emu
机构
*
Department of Computer Science, Lakehead University(湖头大学计算机科学系)
;
Quantum Communications and Computing Research Center and Department of Electrical and Computer Engineering, Memorial University of Newfoundland(新斯科舍纪念大学量子通信与计算研究中心及电气与计算机工程系)
;
Department of Electrical and Computer Engineering, Memorial University of Newfoundland(新斯科舍纪念大学电气与计算机工程系)
专题命中
效率与部署
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI
Bridging the Agent-World Gap: Text World Models for LLM-based Agents
弥合智能体-世界鸿沟:面向基于LLM的智能体的文本世界模型
Yixia Li, Hongru Wang, Peng Lai, Zhiwen Ruan, He Zhu, Youxin Zhu, Ganlong Zhao, Minda Hu, Yun Chen, Sibei Yang, Peng Li, Jeff Z. Pan, Jia Pan, Guanhua Chen, Yang Liu, Guanbin Li
机构
*
Southern University of Science and Technology(南方科技大学)
;
University of Edinburgh(爱丁堡大学)
;
Peking University(北京大学)
;
Sun Yat-sen University(中山大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Shanghai University of Finance and Economics(上海财经大学)
;
Tsinghua University(清华大学)
;
The University of Hong Kong(香港大学)
专题命中
效率与部署
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL
Does the way we write a theory change the program an LLM builds from it? A prospective randomized study of renderer format in LLM theory-to-program translation