arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-11-07 至 2025-11-07 共收录 33 信号源:cs.CL, cs.AI, cs.LG

1. 效率与部署 33 篇

2505.24722 2025-11-07 cs.LG cs.AI 91%

HELM: Hyperbolic Large Language Models via Mixture-of-Curvature Experts

Neil He, Rishabh Anand, Hiren Madhu, Ali Maatouk, Smita Krishnaswamy, Leandros Tassiulas, Menglin Yang, Rex Ying

机构 * Yale University, USA(耶鲁大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract);pretraining(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03830 2025-11-07 cs.CL 90%

Divide, Cache, Conquer: Dichotomic Prompting for Efficient Multi-Label LLM-Based Classification

Mikołaj Langner, Jan Eliasz, Ewa Rudnicka, Jan Kocoń

机构 * Department of Artificial Intelligence, Wroclaw Tech, Poland(人工智能系,沃拉布勒技术学院,波兰)

专题命中 效率与部署 :LLM(title,abstract);prompting(title);large language model(abstract);language model(abstract)

Comments 9 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04002 2025-11-07 cs.LG cs.AI 90%

Memory- and Latency-Constrained Inference of Large Language Models via Adaptive Split Computing

Mingyu Sung, Vikas Palakonda, Suhwan Im, Sunghwan Moon, Il-Min Kim, Sangseok Yun, Jae-Mo Kang

机构 * Department of Artificial Intelligence, Kyungpook National University(人工智能系,庆北国立大学) Department of Electrical and Computer Engineering, Queen’s University(电气与计算机工程系,皇后大学) Department of Information and Communications Engineering, Pukyong National University(信息与通信工程系,浦项国立大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04499 2025-11-07 cs.CL cs.AI 88%

Decoding Emergent Big Five Traits in Large Language Models: Temperature-Dependent Expression and Architectural Clustering

Christos-Nikolaos Zacharopoulos, Revekka Kyriakoglou

机构 * Université Paris 8 Vincennes – Saint-Denis(巴黎第八大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

Comments Accepted at IJCNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09334 2025-11-07 cs.DC 85%

ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments

Youhe Jiang, Fangcheng Fu, Xiaozhe Yao, Taiyi Wang, Bin Cui, Ana Klimovic, Eiko Yoneki

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract)

Comments MLSys 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04036 2025-11-07 cs.AR 85%

PICNIC: Silicon Photonic Interconnected Chiplets with Computational Network and In-memory Computing for LLM Inference Acceleration

Yue Jiet Chong, Yimin Wang, Zhen Wu, Xuanyao Fong

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11916 2025-11-07 cs.DC 85%

Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture

Yu Wu, Tongxuan Liu, Yuting Zeng, Siyu Wu, Jun Xiong, Xianzhe Dong, Hailong Yang, Ke Zhang, Jing Li

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04063 2025-11-07 cs.LG cs.CL 84%

DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization

Yuantian Shao, Yuanteng Chen, Peisong Wang, Jianlin Yu, Jing Lin, Yiwu Yao, Zhihui Wei, Jian Cheng

机构 * Nanjing University of Science and Technology(南京理工大学) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Huawei Technologies Co., Ltd.(华为技术有限公司)

专题命中 效率与部署 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

Comments NeurIPS 2025, 10 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04020 2025-11-07 cs.CL cs.AI 84%

Abductive Inference in Retrieval-Augmented Language Models: Generating and Validating Missing Premises

Shiyin Lin

机构 * Independent Researcher, Mountain View, CA94039, USA(独立研究者)

专题命中 效率与部署 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15482 2025-11-07 cs.CV cs.AI 83%

Comparing Computational Pathology Foundation Models using Representational Similarity Analysis

Vaibhav Mishra, William Lotter

机构 * Dana-Farber Cancer Institute(达纳-法伯癌症研究所) Brigham and Women’s Hospital & Harvard Medical School(布里奇沃特医院及哈佛医学院)

专题命中 效率与部署 :foundation model(title,abstract);language model(abstract);分类 cs.AI

Comments Proceedings of the 5th Machine Learning for Health (ML4H) Symposium

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04477 2025-11-07 cs.DC 82%

Enabling Dynamic Sparsity in Quantized LLM Inference

Rongxiang Wang, Kangyuan Shu, Felix Xiaozhu Lin

专题命中 效率与部署 :LLM(title);large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03844 2025-11-07 cs.MA 82%

ASAP: an Agentic Solution to Auto-optimize Performance of Large-Scale LLM Training

Yuran Ding, Xinwei Chen, Xiaofan Zhang, Zongwei Zhou

专题命中 效率与部署 :LLM(title,abstract);language model(abstract)

Comments This work has been accepted to Workshop on ML for Systems at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03983 2025-11-07 cs.LG math.OC 81%

TwIST: Rigging the Lottery in Transformers with Independent Subnetwork Training

Michael Menezes, Barbara Su, Xinze Feng, Yehya Farhat, Hamza Shili, Anastasios Kyrillidis

专题命中 效率与部署 :LLM(abstract);large language model(abstract);language model(abstract);post-training(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04184 2025-11-07 cs.CL cs.AI 81%

Trustworthy LLM-Mediated Communication: Evaluating Information Fidelity in LLM as a Communicator (LAAC) Framework in Multiple Application Domains

Mohammed Musthafa Rafi, Adarsh Krishnamurthy, Aditya Balu

机构 * Iowa State University(爱荷华州立大学)

专题命中 效率与部署 :LLM(title,abstract);分类 cs.CL、cs.AI

Comments 10 pages, 4 figures. Submitted to IEEE DISTILL 2025 (co-located with IEEE TPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03825 2025-11-07 cs.AI cs.CL cs.CR cs.LG 80%

How Different Tokenization Algorithms Impact LLMs and Transformer Models for Binary Code Analysis

Ahmed Mostafa, Raisul Arefin Nahid, Samuel Mulder

机构 * Ahmed Mostafa, 1Raisul Arefin, Samuel Mulder(未知)

专题命中 效率与部署 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Publication Notice. This paper was published in the BAR 2025 Workshop (with NDSS 2025) and is for research and educational use. Copyright \c{opyright} 2025 Internet Society. All rights reserved. Personal/classroom reproduction is permitted with this notice and full paper citation. All other uses, including commercial, require prior written permission from the Internet Society

Journal ref https://www.ndss-symposium.org/wp-content/uploads/bar2025-final13.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23971 2025-11-07 cs.LG 79%

Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model Training

William Merrill, Shane Arora, Dirk Groeneveld, Hannaneh Hajishirzi

机构 * Allen Institute for AI(人工智能研究所)

专题命中 效率与部署 :language model(title,abstract);分类 cs.LG

Comments Neurips 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17793 2025-11-07 cs.CL 79%

Compression Hacking: A Supplementary Perspective on Informatics Properties of Language Models from Geometric Distortion

Jianxiang Zang, Meiling Ning, Yongda Wei, Shihan Dou, Jiazheng Zhang, Nijia Mo, Binhong Li, Tao Gui, Qi Zhang, Xuanjing Huang

机构 * Computation and Artificial Intelligence Innovative College, Fudan University(复旦大学计算与人工智能创新学院) Beijing University of Posts and Telecommunications(北京邮电大学) Shanghai University of International Business and Economics(上海国际商务经济学院) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 效率与部署 :language model(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16956 2025-11-07 cs.CL 79%

On Multilingual Encoder Language Model Compression for Low-Resource Languages

Daniil Gurgurov, Michal Gregor, Josef van Genabith, Simon Ostermann

机构 * Saarland University(萨尔兰州大学) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心) Kempelen Institute of Intelligent Technologies (KInIT)(Kempelen智能技术研究所) Centre for European Research in Trusted AI (CERTAIN)(可信人工智能欧洲研究中心)

专题命中 效率与部署 :language model(title,abstract);分类 cs.CL

Comments Accepted to SRW AACL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04286 2025-11-07 cs.LG cs.AI 79%

Efficient Reinforcement Learning from Human Feedback via Bayesian Preference Inference

Matteo Cercola, Valeria Capretti, Simone Formentin

专题命中 效率与部署 :LLM(abstract);RLHF(abstract);preference optimization(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04214 2025-11-07 cs.LG cs.CL 79%

Block Rotation is All You Need for MXFP4 Quantization

Yuantian Shao, Peisong Wang, Yuanteng Chen, Chang Xu, Zhihui Wei, Jian Cheng

机构 * Nanjing University of Science and Technology(南京理工大学) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Zhongguancun Academy(中关村学院) School of Computer Science, University of Sydney(悉尼大学计算机科学学院)

专题命中 效率与部署 :large language model(abstract);language model(abstract);post-training(abstract);分类 cs.CL、cs.LG

Comments 9 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22879 2025-11-07 cs.LG cs.AI cs.CL cs.PF 78%

Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models

Hung-Yueh Chiang, Chi-Chih Chang, Natalia Frumkin, Kai-Chiang Wu, Mohamed S. Abdelfattah, Diana Marculescu

专题命中 效率与部署 :post-training(title);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04131 2025-11-07 cs.RO 78%

BFM-Zero: A Promptable Behavioral Foundation Model for Humanoid Control Using Unsupervised Reinforcement Learning

Yitang Li, Zhengyi Luo, Tonghe Zhang, Cunxi Dai, Anssi Kanervisto, Andrea Tirinzoni, Haoyang Weng, Kris Kitani, Mateusz Guzek, Ahmed Touati, Alessandro Lazaric, Matteo Pirotta, Guanya Shi

专题命中 效率与部署 :foundation model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04215 2025-11-07 cs.CR cs.CL 77%

Black-Box Guardrail Reverse-engineering Attack

Hongwei Yao, Yun Xia, Shuo Shao, Haoran Shi, Tong Qiao, Cong Wang

机构 * City University of Hong Kong(香港城市大学) Zhejiang University(浙江大学) Hangzhou Dianzi University(杭州电子科技大学)

专题命中 效率与部署 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22913 2025-11-07 cs.LG 74%

Mustafar: Promoting Unstructured Sparsity for KV Cache Pruning in LLM Inference

Donghyeon Joo, Helya Hosseini, Ramyad Hadidi, Bahar Asgari

机构 * Department of Computer Science, University of Maryland(计算机科学系,马里兰大学) d-Matrix

专题命中 效率与部署 :LLM(title);分类 cs.LG

Comments 20 pages, 9 figures, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07416 2025-11-07 cs.CV cs.CL cs.LG 73%

RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Chest X-ray with Zero-Shot Multi-Task Capability

Jonggwon Park, Byungmu Yoon, Soobum Kim, Kyoyun Choi

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17770 2025-11-07 cs.LG cond-mat.dis-nn 70%

Small Singular Values Matter: A Random Matrix Analysis of Transformer Models

Max Staats, Matthias Thamm, Bernd Rosenow

机构 * Center for Scalable Data Analytics and Artificial Intelligence(可扩展数据与人工智能研究中心) Leipzig University(莱比锡大学) Institute for Theoretical Physics(理论物理研究所)

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.LG

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04172 2025-11-07 cs.IR cs.CL 70%

Transforming Mentorship: An AI Powered Chatbot Approach to University Guidance

Mashrur Rahman, Mantaqa abedin, Monowar Zamil Abir, Faizul Islam Ansari, Adib Reza, Farig Yousuf Sadeque, Niloy Farhan

机构 * Computer Science and Engineering(计算机科学与工程) Brac University(布拉克斯大学)

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.CL

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18948 2025-11-07 cs.LG cs.CC cs.FL 70%

Exact Expressive Power of Transformers with Padding

William Merrill, Ashish Sabharwal

机构 * Allen Institute for AI(艾伦人工智能研究所)

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.LG

Comments Neurips 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10302 2025-11-07 cs.DC 67%

SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference

Liangkun Chen, Zijian Wen, Tian Wu, Xiaoxi Zhang, Chuan Wu

专题命中 效率与部署 :large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18989 2025-11-07 cs.LG cs.AI cs.AR 62%

GENIAL: Generative Design Space Exploration via Network Inversion for Low Power Algorithmic Logic Units

Maxence Bouvier, Ryan Amaudruz, Felix Arnold, Renzo Andri, Lukas Cavigelli

机构 * Huawei, Switzerland(华为,瑞士)

专题命中 效率与部署 :pretraining(abstract);分类 cs.AI、cs.LG

Comments Accepted at the 2026 31st Asia and South Pacific Design Automation Conference (ASP-DAC)

详情

展开后加载摘要…

URL PDF HTML 收藏