arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-12-29 至 2025-12-29 共收录 26 信号源:cs.CL, cs.AI, cs.LG

1. 效率与部署 26 篇

2501.15544 2025-12-29 cs.LG cs.AI 88%

Advancing Generative Artificial Intelligence and Large Language Models for Demand Side Management with Internet of Electric Vehicles

推动生成式人工智能和大语言模型在电动汽车互联网中的需求侧管理

Hanwen Zhang, Ruichen Zhang, Wei Zhang, Dusit Niyato, Yonggang Wen, Chunyan Miao

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出利用生成式人工智能和大语言模型优化电动汽车互联网中的需求侧管理,通过检索增强生成提升能源效率和用户适应性。

Comments 15 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21859 2025-12-29 cs.CL 88%

TimeBill: Time-Budgeted Inference for Large Language Models

TimeBill: 为大语言模型设计的时间预算推理方法

Qi Fan, An Zou, Yehan Ma

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 TimeBill提出了一种时间预算推理框架,通过细粒度响应长度预测和执行时间估计,提升大语言模型在时间敏感任务中的效率和响应性能。

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18564 2025-12-29 cs.AI 85%

Vox Deorum: A Hybrid LLM Architecture for 4X / Grand Strategy Game AI -- Lessons from Civilization V

Vox Deorum:一种用于4X/大战略游戏AI的混合LLM架构——来自《文明V》的启示

John Chen, Sihan Cheng, Can Gurkan, Ryan Lay, Moez Salahuddin

机构 * University of Arizona(亚利桑那大学) Northwestern University(西北大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Independent Researcher(独立研究者)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 Vox Deorum提出了一种混合LLM架构用于4X游戏AI,通过宏观战略推理与子系统协同,展示了LLM在复杂游戏中的应用潜力。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21571 2025-12-29 cs.DC cs.LG 85%

nncase: An End-to-End Compiler for Efficient LLM Deployment on Heterogeneous Storage Architectures

nncase:一种面向异构存储架构高效大语言模型部署的端到端编译器

Hui Guo, Qihang Zheng, Chenghai Huo, Dongliang Guo, Haoqi Yang, Yang Zhang

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 nncase通过端到端编译框架实现高效大语言模型部署,整合三个核心模块提升异构存储架构下的性能表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21352 2025-12-29 cs.SE cs.AI cs.MA 85%

Multi-Agent LLM Committees for Autonomous Software Beta Testing

多智能体大语言模型委员会用于自主软件Beta测试

Sumanth Bharadwaj Hachalli Karanam, Dhiwahar Adhithya Kennady

机构 * New York University(纽约大学)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 多智能体大语言模型委员会通过三轮投票协议实现自主软件Beta测试,显著提高任务成功率和漏洞检测效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05347 2025-12-29 cs.NI 85%

A Queueing Theoretic Perspective on Low-Latency LLM Inference with Variable Token Length

从排队论角度探讨具有可变令牌长度的低延迟LLM推理

Yuqing Yang, Yuedong Xu, Lei Jiao

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract)

AI总结 本文从排队论角度研究了LLM推理中可变令牌长度对低延迟的影响,提出数学模型优化最大令牌限制并分析不同批量策略的排队延迟。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20308 2025-12-29 cs.CL cs.SD eess.AS 83%

SpidR: Learning Fast and Stable Linguistic Units for Spoken Language Models Without Supervision

SpidR:无需监督学习快速稳定的语言单元用于语音语言模型

Maxime Poli, Mahi Luthra, Youssef Benchekroun, Yosuke Higuchi, Martin Gleize, Jiayi Shen, Robin Algayres, Yu-An Chung, Mido Assran, Juan Pino, Emmanuel Dupoux

机构 * ENS-PSL, EHESS, CNRS(ENS-PSL、EHESS、CNRS) FAIR at Meta(Meta的FAIR)

专题命中 效率与部署 :language model(title,abstract);pretraining(abstract);分类 cs.CL

AI总结 SpidR通过自监督学习高效语音表示,提升无监督语音语言建模性能,减少预训练时间

Comments Published in Transactions on Machine Learning Research. 30 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22066 2025-12-29 cs.AR cs.LG cs.PF 77%

Prefill vs. Decode Bottlenecks: SRAM-Frequency Tradeoffs and the Memory-Bandwidth Ceiling

预填与解码瓶颈:SRAM频率权衡及内存带宽天花板

Hannah Atmer, Yuan Yao, Thiemo Voigt, Stefanos Kaxiras

机构 * Uppsala University(乌普萨拉大学)

专题命中 效率与部署 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文研究了SRAM大小和频率对LLM推理能耗与性能的影响,发现高频率和小缓冲区能优化能耗-延迟乘积,平衡低延迟与高能效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10436 2025-12-29 cs.CL 77%

RefactorCoderQA: Benchmarking LLMs for Multi-Domain Coding Question Solutions in Cloud and Edge Deployment

RefactorCoderQA: 评估LLMs在云和边缘部署中多领域编程问题解决方案的基准测试

Shadikur Rahman, Aroosa Hameed, Gautam Srivastava, Syed Muhammad Danish

机构 * York University and Algoma University(约克大学和阿尔戈马大学) Algoma University(阿尔戈马大学) Department of Systems and Computer Engineering, Carleton University(系统与计算机工程系,卡尔顿大学) Department of Math and Computer Science, Brandon University(数学与计算机科学系,布兰登大学)

专题命中 效率与部署 :large language model(abstract);language model(abstract);prompting(abstract);分类 cs.CL

AI总结 RefactorCoderQA通过云-边缘协作架构和RefactorCoder-MoE模型,在多领域编程任务中提升LLMs的性能,准确率达76.84%。

Comments 12 pages, 5 figures, Submitted to IEEE Transactions on Services Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08653 2025-12-29 cs.LG cs.AI cs.CL cs.NE 75%

Accelerating Training Speed of Tiny Recursive Models with Curriculum Guided Adaptive Recursion

通过课程指导自适应递归加速小型递归模型的训练速度

Kaleem Ullah Qasim, Jiashu Zhang

机构 * School of Computing and Artificial Intelligence, Southwest Jiaotong University(计算机与人工智能学院,西南交通大学)

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 CGAR通过课程指导自适应递归方法,提升小型递归模型训练效率,实现加速训练和保持模型质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21008 2025-12-29 cs.CR 75%

GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMs

GateBreaker: 对混合专家LLM的门控引导攻击

Lichao Wu, Sasha Behrouzi, Mohamadreza Rostami, Stjepan Picek, Ahmad-Reza Sadeghi

专题命中 效率与部署 :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 GateBreaker通过门控引导攻击破坏MoE LLM的安全对齐,发现安全机制集中在少量神经元上,禁用这些神经元显著提高攻击成功率。

Comments Accepted by USENIX Security'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14051 2025-12-29 cs.CL cs.AI 73%

GroupDebate: Enhancing the Efficiency of Multi-Agent Debate Using Group Discussion

GroupDebate: 通过群体讨论提升多智能体辩论效率

Tongxuan Liu, Xingyu Wang, Weizhe Huang, Wenjiang Xu, Yuting Zeng, Lei Jiang, Hailong Yang, Jing Li

机构 * University of Science and Technology of China(中国科学技术大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Beihang University(北航)

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出GroupDebate方法,通过将智能体划分为小组并共享中间结果,减少多智能体辩论的token成本,提升效率和准确性。

Comments Accepted by AAMAS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21818 2025-12-29 cs.SE cs.MA 71%

Analyzing Code Injection Attacks on LLM-based Multi-Agent Systems in Software Development

分析基于大语言模型的多智能体系统在软件开发中的代码注入攻击

Brian Bowers, Smita Khapre, Jugal Kalita

专题命中 效率与部署 :LLM(title)

AI总结 本文研究了基于大语言模型的多智能体系统在软件开发中的代码注入攻击问题,提出了一种更健壮的架构,并通过添加安全分析代理提升了系统的安全性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21878 2025-12-29 cs.MA cs.AI 70%

MASFIN: A Multi-Agent System for Decomposed Financial Reasoning and Forecasting

MASFIN:一种用于分解金融推理和预测的多智能体系统

Marc S. Montalvo, Hamed Yaghoobian

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 MASFIN通过整合LLMs与结构化财务指标和非结构化新闻,提出了一种模块化的多智能体系统,以提高金融预测的透明性和可重复性,实现优于基准的短期投资回报。

Comments Accepted to the NeurIPS 2025 Workshop on Generative AI in Finance

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21708 2025-12-29 cs.CL 70%

MoRAgent: Parameter Efficient Agent Tuning with Mixture-of-Roles

MoRAgent: 基于角色混合的参数高效代理调优

Jing Han, Binwei Yan, Tianyu Guo, Zheyuan Bai, Mengyu Zheng, Hanting Chen, Ying Nie

机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 MoRAgent通过角色混合框架实现参数高效代理调优,分解任务为推理、执行和总结三个角色,结合低秩适应技术提升代理性能。

Comments Accepted by ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22009 2025-12-29 cs.CV 67%

iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception

iSHIFT: 轻量级慢-快 GUI 代理与自适应感知

Sarthak Mehrotra, Sairam V C Rebbapragada, Mani Hemanth Reddy Bonthu, Vineeth N Balasubramanian

机构 * Indian Institute of Technology, Bombay(印度理工学院,孟买) Indian Institute of Technology, Hyderabad(印度理工学院,海得拉巴)

专题命中 效率与部署 :large language model(abstract);language model(abstract)

AI总结 iSHIFT是一种轻量级GUI代理,通过慢-快混合推理和灵活标记实现高效与精确的视觉交互。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21720 2025-12-29 cs.LG cs.AI cs.CL cs.IT math.IT 67%

An Information Theoretic Perspective on Agentic System Design

从信息论角度看待代理系统设计

Shizhe He, Avanika Narayan, Ishan S. Khare, Scott W. Linderman, Christopher Ré, Dan Biderman

机构 * Department of Computer Science, Stanford University(斯坦福大学计算机科学系) Department of Statistics, Stanford University(斯坦福大学统计学系) Wu Tsai Neurosciences Institute, Stanford University(斯坦福大学吴泰教授神经科学研究所)

专题命中 效率与部署 :language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 从信息论角度研究代理系统设计,通过互信息估计器评估压缩质量,发现更大压缩器更准确且更高效,可显著降低资源消耗。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19443 2025-12-29 cs.CV 67%

D2Pruner: Debiased Importance and Structural Diversity for MLLM Token Pruning

D2Pruner: 用于MLLM标记剪枝的去偏重要与结构多样性

Evelyn Zhang, Fufu Yu, Aoqi Wu, Zichen Wen, Ke Yan, Shouhong Ding, Biqing Qi, Linfeng Zhang

机构 * Tencent YouTu Lab(腾讯YouTu实验室)

专题命中 效率与部署 :large language model(abstract);language model(abstract)

AI总结 D2Pruner通过结合去偏重要与结构剪枝机制,有效提升MLLM标记剪枝的效率和保真度,尤其在细粒度定位任务中表现突出。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22007 2025-12-29 cs.LG 57%

DuaDeep-SeqAffinity: Dual-Stream Deep Learning Framework for Sequence-Only Antigen-Antibody Affinity Prediction

DuaDeep-SeqAffinity:双流深度学习框架用于仅序列抗原-抗体亲和力预测

Aicha Boutorh, Soumia Bouyahiaoui, Sara Belhadj, Nour El Yakine Guendouz, Manel Kara Laouar

机构 * National School of Artificial Intelligence (ENSIA)(人工智能国家学校)

专题命中 效率与部署 :language model(abstract);分类 cs.LG

AI总结 DuaDeep-SeqAffinity通过双流深度学习框架,仅利用氨基酸序列预测抗原-抗体亲和力,优于现有方法并提高预测精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21985 2025-12-29 cs.CV cs.AI 57%

LVLM-Aided Alignment of Task-Specific Vision Models

通过大视觉语言模型辅助对齐任务特定视觉模型

Alexander Koebler, Lukas Kuhn, Ingo Thon, Florian Buettner

专题命中 效率与部署 :language model(abstract);分类 cs.AI

AI总结 通过大视觉语言模型辅助对齐任务特定视觉模型,提升模型与人类规范的一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21911 2025-12-29 cs.CL 57%

Accelerate Speculative Decoding with Sparse Computation in Verification

通过验证中的稀疏计算加速推测解码

Jikai Wang, Jianchao Tan, Yuxuan Hu, Jiayu Qin, Yerui Sun, Yuchen Xie, Xunliang Cai, Juntao Li, Min Zhang

机构 * Key Laboratory of Data Intelligence and Advanced Computing, Soochow University(数据智能与先进计算 key laboratory,苏州大学)

专题命中 效率与部署 :language model(abstract);分类 cs.CL

AI总结 本文提出了一种稀疏验证框架,通过联合稀疏化注意力、FFN和MoE组件,减少验证阶段的计算成本,并结合检索重用策略提升效率-准确性平衡。

Comments Pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21626 2025-12-29 cs.AI 57%

Multiple-play Stochastic Bandits with Prioritized Arm Capacity Sharing

多玩家随机博弈中的优先级臂容量共享

Hong Xie, Haoran Gu, Yanying Huang, Tao Tan, Defu Lian

专题命中 效率与部署 :LLM(abstract);分类 cs.AI

AI总结 本文提出了一种多玩家随机博弈变种,用于资源分配问题,设计了MSB-PRS-OffOpt算法以优化优先级资源共享机制下的最优播放分配策略。

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21450 2025-12-29 cs.LG 57%

RLLaVA: An RL-central Framework for Language and Vision Assistants

RLLaVA: 一种面向语言和视觉助手的强化学习中心框架

Lei Zhao, Zihao Ma, Boyu Lin, Yuhe Liu, Wenjun Wu, Lei Huang

机构 * SKLCCSE, Institute of Artificial Intelligence, Beihang University, Beijing, China(信息与电子技术学院,人工智能研究院,北京航空航天大学,北京,中国) Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University(未来区块链与隐私计算先进创新中心,北京航空航天大学) Hangzhou International Innovation Institute, Beihang University, Hangzhou, China(杭州国际创新研究院,北京航空航天大学,杭州,中国)

专题命中 效率与部署 :language model(abstract);分类 cs.LG

AI总结 RLLaVA 提出了一种强化学习中心框架,通过解耦算法逻辑与模型架构,实现高效训练和多任务扩展,提升视觉-语言模型性能。

Comments The code is available at https://github.com/TinyLoopX/RLLaVA

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21730 2025-12-29 cs.DC 50%

Hyperion: Low-Latency Ultra-HD Video Analytics via Collaborative Vision Transformer Inference

Hyperion: 通过协作视觉Transformer推理实现低延迟超高清视频分析

Linyi Jiang, Yifei Zhu, Hao Yin, Bo Li

专题命中 效率与部署 :foundation model(abstract)

AI总结 Hyperion通过协作视觉Transformer推理实现低延迟超高清视频分析,提升处理效率和准确性。

Comments Accepted for publication in IEEE INFOCOM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21252 2025-12-29 cs.CV 50%

DreaMontage: Arbitrary Frame-Guided One-Shot Video Generation

DreaMontage:任意帧引导的一次生成视频

Jiawei Liu, Junqiao Li, Jiangfan Deng, Gen Li, Siyu Zhou, Zetao Fang, Shanshan Lao, Zengde Deng, Jianing Zhu, Tingting Ma, Jiayi Li, Yunqiu Wang, Qian He, Xinglong Wu

机构 * Intelligence Creation Team, ByteDance(字节跳动智能创作团队)

专题命中 效率与部署 :SFT(abstract)

AI总结 DreaMontage通过任意帧引导生成技术,实现高效且高质量的一次生成视频,提升电影制作的视觉表现力和内容连贯性。

Comments Project Page: https://dreamontage.github.io/DreaMontage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21508 2025-12-29 cs.CV 50%

Fixed-Budget Parameter-Efficient Training with Frozen Encoders Improves Multimodal Chest X-Ray Classification

固定预算参数高效训练结合冻结编码器提升多模态胸片分类

Md Ashik Khan, Md Nahid Siddique

机构 * Department of Computer Science and Engineering, Indian Institute of Technology Kharagpur, India(计算机科学与工程系,印度理工学院Kharagpur分校) Knight Foundation School of Computing and Information Sciences, Florida International University, Florida, USA(骑士基金会计算与信息科学学院,佛罗里达国际大学)

专题命中 效率与部署 :language model(abstract)

AI总结 本研究通过冻结编码器的参数高效训练策略,在降低计算成本的同时提升了多模态胸片分类的性能。

Comments Accepted at the 2025 28th International Conference on Computer and Information Technology (ICCIT). 6 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏