arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 12294 信号源:cs.CL, cs.AI, cs.LG

1. 效率与部署 2073 篇

2607.15297 2026-07-20 eess.IV cs.MM 新提交 92%

Large Language Model-Enhanced Multi-hop Parallel Image Semantic Communication

大语言模型增强的多跳并行图像语义通信

Bingyan Xie, Jihong Park, Rui Mao, Longyu Zhou, Tianhao Liang, Yongpeng Wu, Wenjun Zhang

专题命中 效率与部署 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract)

AI总结 研究提出LLM-MHPSC框架减轻多跳无线图像传输失真累积,设计粗到细残余压缩方案,开发LLM-RTO并提出自适应跳选择策略,实验表明其优于现有方案,为多跳语义通信应用提供灵活有效解决方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00511 2026-07-02 cs.SE 新提交 92%

Large Language Models for Multi-Lingual Equivalent Mutant Detection: An Extended Empirical Study

大型语言模型用于多语言等价变异检测:一项扩展的实证研究

Honglin Shu, Zhao Tian, Dong Wang, Junji Yu, Jiazhe Zhang, Xuejie Cao, Junjie Chen, Yasutaka Kamei

专题命中 效率与部署 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract)

AI总结 本文首次全面实证研究大型语言模型(LLMs)在等价变异检测(EMD)中的应用,使用Java和C变异对评估其性能,发现基于LLM的方法在F1分数上优于传统方法,且微调后的代码嵌入策略检测精度最高,同时展现出跨语言泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24569 2026-08-26 cs.AI cs.MA 新提交 92%

When "Must" Becomes "Maybe": Constraint Weakening in LLM Agent Workflows

当“必须”变为“可能”:LLM智能体工作流中的约束弱化

Yiheng Sun, Huifei Wang, Yancheng Zhu, Zhenyu Li, Zebin Zhao, Yifan Yuan

机构 * Shenzhen University(深圳大学)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 该研究针对LLM智能体工作流中交接转换弱化状态约束的问题,通过1296个受控合成场景实验,发现常规交接压缩会导致高失活和禁止动作率,恢复全部状态字段可解决该问题。

Comments 21 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16477 2026-08-18 cs.LG 新提交 92%

Pallas: A Proactive KV Cache Migration Framework for LLM Inference in AI-RAN

Pallas:面向AI-RAN中LLM推理的主动KV缓存迁移框架

Tianhang Ding, Jianchun Liu, Hongli Xu

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 Pallas是面向AI-RAN中LLM推理的主动KV缓存迁移框架,通过切换前在目标基站预准备推理状态,结合在线调度器优化,显著降低服务中断时间与令牌间延迟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16201 2026-08-18 cs.LG 新提交 92%

Multi-Granularity Sentiment Integration for LLM-Based Multimodal Sentiment Analysis

面向基于大语言模型(LLM)的多模态情感分析的多粒度情感集成

Shanshan Lin, Yuesheng Wu, Chao Chen, Yizhe Yang, Zhihao Chen, Zexian Yang, Xiangwen Liao

机构 * Fuzhou University(福州大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) Jiangxia University(江夏大学)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 该研究提出MGSI多粒度情感集成框架,通过多尺度编码、文本引导对齐等优化,提升基于LLM的多模态情感分析性能,在四个公开基准上效果优于冻结LLM基线。

Comments Accepted to NLPCC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13573 2026-08-17 cs.AI 新提交 92%

A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing

LLM 服务的一年:工作负载演变、缓存与负载均衡

William Nixon, Jon Durbin, Florian Standhartinger, Haryadi S. Gunawi, Juncheng Yang

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 该研究通过分析 Chutes 一年的 LLM 生产请求轨迹,揭示了 LLM 服务工作负载的演变规律与用户-模型结构,并将发布完整轨迹供后续研究使用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10532 2026-08-12 cs.NI cs.LG 新提交 92%

Benchmarking LLM-Guided Control-Plane Policies for Backend Fault Isolation in HAProxy

对HAProxy中后端故障隔离的LLM引导控制平面策略进行基准测试

Aman Chauhan, Vishnu Pendyala

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文测试了不同规模和架构的LLM在HAProxy后端故障隔离中的策略效果,发现约3B有效参数的能力阈值,达标后可显著降低5xx错误,但会增加延迟与令牌开销,高效方案为阈值以上模型的非推理模式。

Comments 43 pages, 6 figures, 15 tables. Submitted to Journal of Network and Computer Applications (Elsevier)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09949 2026-08-12 cs.AI physics.bio-ph 新提交 92%

Closed-Loop LLM Co-Pilots for Digital Agriculture

面向数字农业的闭环大语言模型(LLM)协作副驾驶系统

Serge Kernbach

机构 * CYBRES GmbH(CYBRES有限公司)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本研究开发闭环LLM协作副驾驶系统,基于49通道植物传感器网络实现数字农业自主控制,通过三案例验证可缩短生产周期、降低能耗,提升成本价值比。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06867 2026-08-10 cs.CL 新提交 92%

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers

LLMRouter:用于开发、评估和部署LLM路由器的统一基础设施

Tao Feng, Fangxu Yu, Haozhen Zhang, Zhongjie Dai, Liangqi Yuan, Zijie Lei, Weizhi Zhang, Kunlun Zhu, Haodong Yue, Keyang Xuan, Ge Liu, Jiaxuan You

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Maryland, College Park(马里兰大学帕克分校) Nanyang Technological University(南洋理工大学) Purdue University(普渡大学) University of Illinois Chicago(芝加哥伊利诺伊大学)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 针对现有LLM路由器难以公平比较和扩展的问题,研究提出LLMRouter统一基础设施,构建含多类任务的基准xRouteBench,发现学习型路由器等的性能优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06557 2026-08-10 cs.DC cs.LG 新提交 92%

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving

Cascade:利用感知SLO的延迟预算实现公平且高吞吐量的LLM推理服务

Muhammad Adnan, Rohan Mahapatra, Prashant J. Nair, Daniel Berger, Pantea Zardoshti, Rodrigo Fonseca, Esha Choukse

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 Cascade是一款LLM服务系统,利用单请求延迟预算协同调度与KV缓存管理,在保留异构请求公平性的同时,使LLM推理服务吞吐量最高提升2.4倍,SLO违规率降低40%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02995 2026-08-05 cs.CR cs.AI 新提交 92%

SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems via Deterministic Side Channels

SparSEEty:通过确定性侧信道从利用稀疏性的大语言模型(LLM)服务系统中提取 token

Yongwan Jo, Jinyoung Park, Euihyun Lee, Dokyung Song

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 SparSEEty 是一种针对利用稀疏性的 LLM 服务系统的 token 提取攻击,通过 CVM 侧信道构建预言机、降低监控开销并反转激活轨迹,可高准确率重建 token,监控开销仅 3.7%-7.2%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26571 2026-07-30 cs.LG cs.SE 新提交 92%

From Tokens to Watt-hours: Analytical Energy Estimation for LLM Inference on Modern GPUs

从Token到瓦时:现代GPU上大语言模型(LLM)推理的分析式能耗估算

Tina Vartziotis, Rodopi Kosteli, Elli Vartziotis, George Dasoulas, Michael Keckeisen, Konstantinos Skianis, Sotirios Kotsopoulos, Francesca Dominici

机构 * National Technical University of Athens(雅典国家技术大学) Harvard University(哈佛大学) TWT GmbH Science & Innovation(TWT有限责任公司科学与创新部) NIKI Ltd Digital Engineering(尼基数字工程有限公司) University of Ioannina(约阿尼纳大学) National and Kapodistrian University of Athens(雅典国立卡波迪斯特里亚大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 该研究提出一种无需直接测量的GPU级LLM推理能耗分析估算方法,适用于模型比较、绿色编码分析及设计时评估。

Comments 20 pages, 3 figures, 6 tables. Accepted for oral presentation at the GREEN-AI Workshop, co-located with ECML-PKDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23370 2026-07-03 cs.CR cs.LG cs.OS 新提交 92%

FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation

FlexServe: 一种面向移动设备的快速安全LLM服务系统,具有灵活的资源隔离

Yinpeng Wu, Yitong Chen, Lixiang Wang, Jinyu Gu, Zhichao Hua, Yubin Xia

机构 * Institute of Parallel and Distributed Systems(并行与分布式系统研究所)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 针对移动设备上LLM推理中TrustZone资源隔离不灵活和安全管理低效的问题,提出FlexServe系统,通过解耦安全资源的访问权限与管理权限,实现快速安全的推理,平均TTFT加速比达10.05倍。

Comments Repeated paper uploading due to mistakes. See arXiv:2603.09046

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26666 2026-07-02 cs.LG 新提交 92%

PersistentKV: Page-Aware Decode Scheduling for Long-Context LLM Serving on Commodity GPUs

PersistentKV: 面向商用GPU上长上下文LLM服务的页面感知解码调度

Muhammad Ahmed

机构 * Muhammad Ahmed(穆罕默德·阿赫姆)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 针对长上下文LLM服务中KV缓存移动成为瓶颈的问题,提出PersistentKV引擎,通过页面感知调度和分组查询注意力优化解码,在RTX 3060上实现1.06-1.40倍吞吐提升。

Comments 7 pages, 3 tables; workshop paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31635 2026-07-01 eess.SY cs.AI cs.MA cs.SY 新提交 92%

A Tutorial on Autonomous Fault-Tolerant Control Using Knowledge-Grounded LLM Agents

基于知识接地LLM代理的自主容错控制教程

Javal Vyas, Milapji Singh Gill, Artan Markaj, Felix Gehlhoff, Mehmet Mercangöz

机构 * Autonomous Industrial Systems Laboratory, Imperial College London(帝国理工学院自主工业系统实验室) Institute of Automation Technology, Helmut Schmidt University(赫尔穆特·施密特大学自动化技术研究所)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出利用LLM代理作为约束监督规划器,通过外部验证器确保安全,辅助过程工厂故障恢复决策,并提供了可执行的Python环境。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22327 2026-06-25 cs.AI 新提交 92%

Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice

面向LLM服务的几何感知在线调度:从理论界到系统实践

Li Kong, Qi Qi, Yinyu Ye, Zijie Zhou

机构 * Gaoling School of Artificial Intelligence(人工智能学院) Renmin University of China(中国人民大学) Department of Management Science and Engineering(管理科学与工程系) Stanford University(斯坦福大学) Department of Industrial Engineering and Decision Analytics(工业工程与决策分析系) HKUST(香港科技大学)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 针对LLM推理中KV缓存动态内存占用问题,提出Smallest Volume First (SVF)算法,通过几何感知调度优化性能,理论证明将竞争比从48降至5,并在vLLM中实现即插即用,显著降低平均和尾部延迟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22179 2026-06-23 cs.CL 新提交 92%

The Score Granularity Gap in Black-Box LLM Classification: A Comparative Study of Confidence Constructions

黑盒LLM分类中的分数粒度差距:置信度构建的比较研究

Ao Sun, Tian Sun, Jiaxing Geng

机构 * institutetext: 1 1(机构1) institutetext: 2 2(机构2) institutetext: 3 3(机构3)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 研究黑盒LLM分类中置信度分数的粒度差距,比较七种构建方法,发现单次口头置信度排名良好但取值有限,多查询聚合可帮助弱模型但可能降低强模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18272 2026-06-19 cs.NI cs.AI cs.SY eess.SY 新提交 92%

Mitigating Anchoring Bias in LLM-Based Agents for Energy-Efficient 6G Autonomous Networks

缓解基于LLM的智能体在节能6G自主网络中的锚定偏差

Hatim Chergui, Claudia Carballo González, Farhad Rezazadeh, Merouane Debbah

机构 * i2CAT Foundation(i2CAT基金会) Universitat Politècnica de Catalunya(政治技术大学) Research Institute for Digital Future(数字未来研究院)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出一种基于截断三参数威布尔分布的随机锚定策略,缓解LLM智能体在6G网络切片中的锚定偏差,结合CVaR数字孪生保障SLA尾延迟,实现高达25%的节能。

Comments 7 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17683 2026-06-17 cs.CL cs.PL 新提交 92%

Bridging Functional Correctness and Runtime Efficiency Gaps in LLM-Based Code Translation

弥合基于LLM的代码翻译中的功能正确性与运行时效率差距

Longhui Zhang, Jiahao Wang, Chenhao Hu, Bingyu Liang, Jing Li, Min Zhang

机构 * Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学(深圳))

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 提出SwiftTrans框架,通过多视角探索和差异感知选择两阶段,利用并行上下文学习和差异比较,同时提升LLM代码翻译的正确性和运行时效率。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15508 2026-06-16 cs.AI 新提交 92%

ToolMenuBench: Benchmarking Tool-Menu Filtering Strategies for Reliable and Efficient LLM Agents

ToolMenuBench: 用于可靠高效LLM智能体的工具菜单过滤策略基准测试

Rahul Suresh Babu, Laxmipriya Ganesh Iyer

机构 * Independent Researcher(独立研究者) United States of America(美国)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出ToolMenuBench基准,评估多步LLM智能体中工具菜单构建策略,发现因果最小工具过滤在任务成功率、令牌使用和风险暴露间取得最佳平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11244 2026-06-11 cs.AR cs.AI 新提交 92%

SPEAR: A System for Post-Quantization Error-Adaptive Recovery Enabling Efficient Low-Bit LLM Serving

SPEAR: 一种后量化误差自适应恢复系统,实现高效低比特LLM服务

Hongyuan Liu, Yawei Li, Zhiqiang Que, Qinli Yang, Junming Shao, Guosheng Hu

机构 * University of Electronic Science and Technology of China(电子科学与技术大学) University of Bristol(布里斯托大学) ETH Zurich(苏黎世联邦理工学院)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 针对低比特量化导致LLM质量下降的问题,提出SPEAR系统,通过输入感知的门控误差补偿器(EC)选择性修正高误差层,结合自适应内核融合调度和SLO感知调度器,在<1%内存开销下恢复W4与FP16之间56-75%的困惑度差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10286 2026-06-10 cs.AI 新提交 92%

Sim2Schedule: A Simulator-Guided LLM Framework for Autonomous Open-Pit Mine Scheduling

Sim2Schedule: 一种模拟器引导的LLM框架用于自主露天矿调度

Mustavi Ibne Masum, Thiago Eustaquio Alves de Oliveira, Mahzabeen Emu

机构 * Department of Computer Science, Lakehead University(湖头大学计算机科学系) Quantum Communications and Computing Research Center and Department of Electrical and Computer Engineering, Memorial University of Newfoundland(新斯科舍纪念大学量子通信与计算研究中心及电气与计算机工程系) Department of Electrical and Computer Engineering, Memorial University of Newfoundland(新斯科舍纪念大学电气与计算机工程系)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出模拟器引导的LLM框架,将地质约束编码到动作生成中,零样本生成可解释调度方案,在保持线性计算时间下恢复MILP最优NPV的94%-99%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09549 2026-06-09 cs.CR cs.AI 新提交 92%

SecureClaw: Clawing Back Control of LLM Agents

SecureClaw: 夺回对LLM智能体的控制

Yuhan Ma, Stefan Schmid

机构 * TU Berlin(柏林技术大学)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 针对工具使用型LLM智能体的双重安全漏洞,提出双边界架构SecureClaw,在效果汇点实施授权、在读边界实施明文隔离,通过预览-提交协议和可信网关实现安全控制,在多个基准上保持可用性的同时将攻击成功率降至接近零。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09388 2026-06-09 cs.LG 新提交 92%

Distilling Safe LLM Systems via Soft Prompts for On Device Settings

通过软提示蒸馏安全的设备端LLM系统

Motasem Alfarra, Cristina Pinneri, Dana Kianfar, Mohammed Almousa, Christos Louizos

机构 * Qualcomm AI Research(高通人工智能研究院)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 针对资源受限设备上部署安全大语言模型(LLM)的挑战,提出基于软提示与蒸馏训练的安全对齐方法,在最小化额外计算开销的同时实现优越的安全-有用性权衡。

Comments Accepted to UAI 2026

Journal ref 42nd Conference on Uncertainty in Artificial Intelligence 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09032 2026-06-09 cs.CL 新提交 92%

Bridging the Agent-World Gap: Text World Models for LLM-based Agents

弥合智能体-世界鸿沟:面向基于LLM的智能体的文本世界模型

Yixia Li, Hongru Wang, Peng Lai, Zhiwen Ruan, He Zhu, Youxin Zhu, Ganlong Zhao, Minda Hu, Yun Chen, Sibei Yang, Peng Li, Jeff Z. Pan, Jia Pan, Guanhua Chen, Yang Liu, Guanbin Li

机构 * Southern University of Science and Technology(南方科技大学) University of Edinburgh(爱丁堡大学) Peking University(北京大学) Sun Yat-sen University(中山大学) The Chinese University of Hong Kong(香港中文大学) Shanghai University of Finance and Economics(上海财经大学) Tsinghua University(清华大学) The University of Hong Kong(香港大学)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文系统综述了面向基于LLM的智能体的文本世界模型,围绕形式化框架和智能体生命周期,涵盖基础定义、构建范式、应用(训练时经验合成与推理时规划、验证、适应)及评估,旨在整合该领域并明确设计空间与开放挑战。

Comments Code: https://github.com/sustech-nlp/awesome-text-world-models

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23814 2026-08-26 cs.LG cs.AI cs.CL 新提交 91%

Learning to Grade Efficiently: A Bandit-Driven Prompt-Selection Framework for Low-Cost LLM Essay Scoring

高效评分学习:面向低成本大语言模型作文评分的多臂老虎机驱动提示选择框架

Olga Manakina, Igor Bogdanov

机构 * Carleton University(卡尔顿大学)

专题命中 效率与部署 :LLM(title,summary_cn);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本研究针对自动作文评分的固定提示选择策略成本高的问题,提出多臂老虎机驱动的自适应提示选择框架,在保证准确率的同时减少78.4%的LLM调用,生成成本-可靠性学习曲线,为教育技术平台提供平衡成本与有效性的方案。

Comments Accepted as a presentation at the EDM 2025 Workshop on Educational Data Mining in Writing and Literacy Instruction

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14376 2026-08-17 cs.OS cs.PF 新提交 91%

CoRun: Padding is Simple and Efficient for Deterministic LLM Inference

CoRun:填充对于确定性大语言模型(LLM)推理而言简单且高效

Shiju Zhao, Jiacheng Yang, Qihang Chen, Junhao Hu, Jiaqi Zheng, Guihai Chen, Xusheng Chen

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract)

AI总结 CoRun利用LLM内核的位置不变性,通过独立预填充和固定形状批量解码的调度系统,在保证确定性推理的同时提升了LLM的吞吐量并降低了延迟。

Comments 13 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10314 2026-08-12 cs.SE physics.comp-ph 新提交 91%

Does the way we write a theory change the program an LLM builds from it? A prospective randomized study of renderer format in LLM theory-to-program translation

我们撰写理论的方式会改变大语言模型(LLM)据此构建的程序吗?一项针对LLM理论转程序翻译中渲染器格式的前瞻性随机研究

Andre Panossian

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract)

AI总结 该研究通过前瞻性随机实验发现,LLM理论转程序翻译中渲染器格式未产生预测的行为几何,为该领域规范格式效应设定了边界并提供了可审计研究工具。

Comments 112 pages, 2 figures; includes supplementary material and five reproducibility annexes. Data: https://doi.org/10.5281/zenodo.21876518. Software: https://doi.org/10.5281/zenodo.21879576

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02244 2026-08-04 cs.DC cs.SY eess.SY 新提交 91%

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling

通过资源公平调度实现批量大语言模型(LLM)服务的效率与成本对齐

Dayi Yao, Zijie Zhou

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract)

AI总结 本文针对批量LLM服务中异构请求导致的资源分配低效问题,提出ISJL算法,实现高吞吐量的同时对齐最大驱动批次成本与token计量收入,达到3/4的竞争比下界,平衡了FCFS与LJF的优缺点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01250 2026-08-04 eess.SY cs.SY 新提交 91%

Smoothing the Ramp, Not the Peak: Scheduling-Induced Power Dynamics of LLM Inference and Their Grid-Scale Consequences

平滑斜坡而非峰值:LLM推理的调度诱导功率动态及其电网级后果

Pan Li, Yize Chen, Xia Miao, Dai Wang

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract)

AI总结 该研究揭示LLM分块预填充调度可降低功率斜坡速率,随系统饱和度提升效益增大,经转化为电网调节备用采购问题后,可减少电网快速斜坡备用容量,为运营商提供无成本需求侧塑形工具。

Comments 15 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏