arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 22280 信号源:cs.CL, cs.AI, cs.LG

1. 效率与部署 22280 篇

2305.16617 2024-06-05 cs.LG cs.AI cs.CL 87%

Efficient Detection of LLM-generated Texts with a Bayesian Surrogate Model

Yibo Miao, Hongcheng Gao, Hao Zhang, Zhijie Deng

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.05465 2024-05-22 cs.LG cs.AI cs.CL 87%

Vidur: A Large-Scale Simulation Framework For LLM Inference

Amey Agrawal, Nitin Kedia, Jayashree Mohan, Ashish Panwar, Nipun Kwatra, Bhargav Gulavani, Ramachandran Ramjee, Alexey Tumanov

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17546 2024-05-01 cs.LG cs.AI cs.CL stat.ML 87%

Probabilistic Inference in Language Models via Twisted Sequential Monte Carlo

Stephen Zhao, Rob Brekelmans, Alireza Makhzani, Roger Grosse

专题命中 效率与部署 :language model(title,abstract);large language model(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13835 2024-03-22 cs.LG cs.AI cs.CL cs.DB 87%

SMART: Automatically Scaling Down Language Models with Accuracy Guarantees for Reduced Processing Fees

Saehan Jo, Immanuel Trummer

专题命中 效率与部署 :language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.19446 2024-03-01 cs.LG cs.AI cs.CL 87%

ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Yifei Zhou, Andrea Zanette, Jiayi Pan, Sergey Levine, Aviral Kumar

专题命中 效率与部署 :language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05605 2024-03-01 cs.CL cs.AI cs.LG 87%

Memory Injections: Correcting Multi-Hop Reasoning Failures during Inference in Transformer-Based Language Models

Mansi Sakarvadia, Aswathy Ajith, Arham Khan, Daniel Grzenda, Nathaniel Hudson, André Bauer, Kyle Chard, Ian Foster

专题命中 效率与部署 :language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Oral Presentation at BlackboxNLP Workshop at EMNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11960 2024-02-20 cs.LG cs.AI cs.CL 87%

DB-LLM: Accurate Dual-Binarization for Efficient LLMs

Hong Chen, Chengtao Lv, Liang Ding, Haotong Qin, Xiabin Zhou, Yifu Ding, Xuebo Liu, Min Zhang, Jinyang Guo, Xianglong Liu, Dacheng Tao

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10076 2024-02-16 cs.LG cs.AI cs.CL 87%

QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference

Taesu Kim, Jongho Lee, Daehyun Ahn, Sarang Kim, Jiwoong Choi, Minkyu Kim, Hyungjun Kim

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 9 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.12028 2023-12-21 cs.SE 87%

Prompt Sapper: A LLM-Empowered Production Tool for Building AI Chains

Yu Cheng, Jieshan Chen, Qing Huang, Zhenchang Xing, Xiwei Xu, Qinghua Lu

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);foundation model(abstract)

Comments 23 pages, 5 figures, accepted to TOSEM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.00502 2023-12-08 cs.LG cs.AI cs.CL 87%

Efficient LLM Inference on CPUs

Haihao Shen, Hanwen Chang, Bo Dong, Yu Luo, Hengyu Meng

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments NeurIPS'2023 on Efficient Natural Language and Speech Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00960 2023-12-05 cs.CL cs.AI cs.LG 87%

The Cost of Compression: Investigating the Impact of Compression on Parametric Knowledge in Language Models

Satya Sai Srinath Namburi, Makesh Sreedhar, Srinath Srinivasan, Frederic Sala

专题命中 效率与部署 :language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to EMNLP 2023 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.03799 2023-11-08 cs.CV 87%

Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation Models

Yichao Cao, Qingfei Tang, Xiu Su, Chen Song, Shan You, Xiaobo Lu, Chang Xu

专题命中 效率与部署 :foundation model(title,abstract);LLM(abstract);large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01434 2023-10-04 cs.CL cs.AI cs.LG 87%

Revolutionizing Mobile Interaction: Enabling a 3 Billion Parameter GPT LLM on Mobile

Samuel Carreira, Tomás Marques, José Ribeiro, Carlos Grilo

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25203 2026-05-26 cs.LG cs.AI cs.LO 87%

Influence-Inspired Spectral Rotations for Extreme Low-Bit LLM Quantization

基于影响启发的谱旋转用于极端低位LLM量化

Gorgi Pavlov

机构 * Lehigh University(莱斯大学)

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.AI、cs.LG

AI总结 本文利用伴随理论论文的影响自适应Walsh几何,通过WHT旋转和列缩放结合重构误差量化器,实现极端低位权重量化,在多个模型上降低困惑度15-58%。

Comments 14 pages, no figures. Companion application paper to arXiv:2605.01637 (theory). Code and pinned eval stack: https://github.com/gogipav14/spectral-llm

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07985 2026-04-07 cs.LG cs.AI cs.CR 87%

Fewer Weights, More Problems: A Practical Attack on LLM Pruning

更少的权重,更多的问题:对LLM剪枝的一种实用攻击

Kazuki Egashira, Robin Staab, Thibaud Gloaguen, Mark Vero, Martin Vechev

机构 * ETH Zurich(苏黎世联邦理工学院)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文首次揭示了现代LLM剪枝方法可能被恶意利用的问题,通过构造看似无害的模型,在剪枝后表现出恶意行为,展示了剪枝过程中的安全风险。

Comments ICLR 2026. Code: https://github.com/eth-sri/llm-pruning-attack

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18638 2025-11-14 cs.CR cs.AI cs.CL 87%

Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation

Daniel Schwartz, Dmitriy Bespalov, Zhe Wang, Ninad Kulkarni, Yanjun Qi

机构 * Amazon Bedrock Science(亚马逊Bedrock科学) Drexel University(德雷塞尔大学) University of Virginia(弗吉尼亚大学)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments 14 pages, 5 figures; published in EMNLP 2025 ; Code at: https://github.com/dsbuddy/GAP-LLM-Safety

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01698 2025-05-16 cs.AR cs.AI cs.DC cs.LG 87%

Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models

Abhimanyu Bambhaniya, Ritik Raj, Geonhwa Jeong, Souvik Kundu, Sudarshan Srinivasan, Suvinay Subramanian, Midhilesh Elavazhagan, Madhu Kumar, Tushar Krishna

机构 * Georgia Institute of Technology(佐治亚理工学院) Meta Intel Labs(英特尔实验室) Intel(英特尔) Google(谷歌)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

Comments 19 Pages, https://github.com/abhibambhaniya/GenZ-LLM-Analyzer, https://genz-llm-analyzer.streamlit.app/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09142 2025-05-15 cs.DC cs.AI cs.LG 87%

ELIS: Efficient LLM Iterative Scheduling System with Response Length Predictor

Seungbeom Choi, Jeonghoe Goo, Eunjoo Jeon, Mingyu Yang, Minsung Jang

机构 * Cloud Research Team, Samsung SDS(三星SDS云研究团队)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

Comments 13 pages, 5 figures. Cloud-native LLM scheduling system with latency-aware inference optimization

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15204 2025-02-11 cs.CL cs.AI cs.HC 87%

Can Unconfident LLM Annotations Be Used for Confident Conclusions?

Kristina Gligorić, Tijana Zrnic, Cinoo Lee, Emmanuel J. Candès, Dan Jurafsky

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments Please cite as: Can Unconfident LLM Annotations Be Used for Confident Conclusions? Kristina Gligorić, Tijana Zrnic, Cinoo Lee, Emmanuel Candès, and Dan Jurafsky. NAACL, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02721 2024-10-15 cs.CL cs.AI 87%

Self-Control of LLM Behaviors by Compressing Suffix Gradient into Prefix Controller

Min Cai, Yuchen Zhang, Shichang Zhang, Fan Yin, Dan Zhang, Difan Zou, Yisong Yue, Ziniu Hu

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments Website: https://llm-self-control.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16434 2024-06-21 cs.AI cs.CL cs.NE 87%

The Importance of Directional Feedback for LLM-based Optimizers

Allen Nie, Ching-An Cheng, Andrey Kolobov, Adith Swaminathan

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments Accepted and Presented at Foundation Models for Decision Making at NeurIPS 2023 (December 15, 2023). Work completed from June 2023 to September 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18933 2026-08-20 cs.SE cs.AI 新提交 86%

SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution

SkillForge:面向项目特定问题解决的自蒸馏智能体

Silin Chen, Han Li, Xiaodong Gu, Yuling Shi, Haibing Guan

专题命中 效率与部署 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 SkillForge是一种主动从代码库获取项目特定知识的自蒸馏框架,通过合成问题蒸馏技能,可提升LLM智能体解决特定代码库软件问题的性能。

Comments Our code and data are available at https://github.com/cslsolow/SkillForge

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18884 2026-08-20 cs.AI 新提交 86%

Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models

面向大语言模型的无训练推理时自我反思与成本受限早停机制

Wei Yu, Suxing Liu, Minjie Yu, Jiahao Wang, Zhijian Zheng, Haocheng Deng, Bing Li

机构 * School of Digital Arts, Jiangxi Arts & Ceramics Technology Institute(江西工艺美术职业技术学院数字艺术学院) School of Computing, Universiti Sains Malaysia(马来西亚理科大学计算机学院)

专题命中 效率与部署 :large language model(title);language model(title);LLM(abstract);分类 cs.AI

AI总结 本文提出无训练的推理时协议EvoResearcher,通过成本受限的自我反思实现大语言模型早停,在多个推理基准上验证其成本受限自我验证的价值。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18263 2026-08-20 cs.LG stat.ML 新提交 86%

SIGMA: Symmetry-aware, Intelligent, Geometric, Multi-objective Adaptive Control for Robust, Dependable Traffic Management

SIGMA:面向鲁棒可靠交通管理的感知对称性、智能型、几何化多目标自适应控制

Pratham Payra, Jagadish B, Tanmay Sen, Tanujit Chakraborty

机构 * Indian Statistical Institute(印度统计研究所) Sorbonne University Abu Dhabi(阿布扎比索邦大学) Sorbonne Center for Artificial Intelligence(索邦人工智能中心) SQC & OR(统计质量控制与运筹学研究室)

专题命中 效率与部署 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文提出SIGMA框架,结合LLM实现自适应交通信号多目标控制,在SUMO仿真中较基准控制器降低等待与排队时长、提升通行量,且具备鲁棒性与统计可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15016 2026-08-18 cs.CR cs.AI 新提交 86%

Hierarchical Agentic Incident Response with Digital-Twin-Validated Attack Inference

结合数字孪生验证攻击推理的分层智能体事件响应

Yiran Gao, Juntao Chen, Tao Li

专题命中 效率与部署 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 该研究针对网络事件响应效率低的问题,提出整合LLM攻击推理、滚动规划与数字孪生验证的分层智能体响应框架,在33组件企业网络测试台的多阶段攻击场景中,其恢复成功率较前沿LLM基线提升18%-31%。

Comments 2026 IEEE Conference on Communications and Network Security

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07672 2026-08-18 cs.CR cs.CL 版本更新 86%

DYNASHIELD: A Black-Box Moving Target Defense for LLMs via Dynamic Decoding Customization

DYNASHIELD:一种通过动态解码定制实现的大语言模型黑盒移动目标防御

Xiaoqun Liu, Weiming Qi, Qiben Yan

专题命中 效率与部署 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 DYNASHIELD是一种无需访问模型内部或重新训练的黑盒LLM防御框架,通过动态定制解码配置降低了4种越狱攻击的成功率,同时保持响应质量且开销极小。

Comments Accepted by The 29th International Symposium on Research in Attacks, Intrusions and Defenses (RAID) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13576 2026-08-17 cs.HC cs.LG q-bio.NC 新提交 86%

BCIJelly: An integrated ecosystem for brain-computer interface research

BCIJelly:用于脑机接口研究的集成生态系统

Liyuan Han, Xinrui Yang, Tianyu Zheng, Qizhi Yang, Yitao Qin, Liang Chen, Qinglai Wei, Binjie Hong, Xinhe Zhang, Rui Xiong, Yong Gu, Mu-ming Poo, Bo Xu, Chengyu Li, Tielin Zhang

专题命中 效率与部署 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 BCIJelly是集成多类BCI数据集、解码器及工具的Python生态系统,可通过AAS与LLM支持多任务跨物种解码,经多范式多物种验证,为BCI研究提供统一可扩展的开发部署基础设施。

Comments 67 pages, 6 figures, 7 extended data figures, 20 supplementary tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10538 2026-08-17 cs.AI 版本更新 86%

SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models

SKILLER:面向小型语言模型可复用技能提取的语言级强化学习框架

Chenhao Dang, Siyuan Xiong, Conghui He, Weijia Li

专题命中 效率与部署 :language model(title,abstract);small language model(title);分类 cs.AI

AI总结 研究针对小型语言模型技能生成成本高的问题,提出SKILLER框架,经实验在多基准上优于现有方法,性能接近闭源模型,可大幅降低智能体技能部署成本。

Comments 15 pages, 6 figures, and 8 tables. Project and code: https://github.com/DANG-ai/SKILLER

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08756 2026-08-14 cs.AI cs.NE 版本更新 86%

AHD Agent: Agentic Reinforcement Learning for Automatic Heuristic Design

AHD代理:用于自动启发式设计的代理强化学习

Haoze Lv, Ning Lu, Ziang Zhou, Yew-Soon Ong, Shengcai Liu

机构 * Guangdong Provincial Key Laboratory of Brain-Inspired Intelligent Computation(广东脑启发智能计算重点实验室) Department of Computer Science and Engineering, Southern University of Science and Technology(南方科技大学计算机科学与工程系) Zhongguancun Academy(中关村学院) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 效率与部署 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出AHD Agent,一种多轮框架,使LLM能主动决定生成启发式或调用工具获取环境证据,通过代理强化学习提升自动启发式设计效率。

Comments 8 pages, 7 figures for main content

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08713 2026-08-11 cs.CV cs.AI 新提交 86%

Resolution Meets Reduction: Efficient Visual Context for 3D Radiology Report Generation

分辨率与压缩结合:面向3D放射报告生成的高效视觉上下文

Jonathan Suprijadi, Raphael Stock, Moritz Langenberg, David Zimmerer, Kim-Celine Kahl, Stefan Denner, Yannick Kirchhoff, Karol Gotkowski, Maximilian Rokuss, Jeremias Traub, Tassilo Wald, Constantin Ulrich, Klaus Maier-Hein

机构 * German Cancer Research Center (DKFZ)(德国癌症研究中心)

专题命中 效率与部署 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 该研究针对3D放射报告生成中视觉序列的计算瓶颈,通过实验评估不同视觉编码器、投影器和LLM,提出解剖学引导的ROI裁剪等策略,在两个数据集上取得最先进的临床macro F1结果。

详情

展开后加载摘要…

URL PDF HTML 收藏