arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-01-23 至 2026-01-23 共收录 170 信号源:cs.CL, cs.AI, cs.LG

1. 效率与部署 41 篇

2601.15482 2026-01-23 cs.LG cs.AI 86%

Martingale Foresight Sampling: A Principled Approach to Inference-Time LLM Decoding

鞅前瞻性采样:一种推断时间LLM解码的原理性方法

Huayu Li, ZhengXiao He, Siyuan Tian, Jinghao Wen, Ao Li

机构 * University of Arizona(亚利桑那大学) Microsoft Research Asia(微软亚洲研究院) Villanova University(维拉诺瓦大学)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出鞅前瞻性采样,通过概率论原理改进LLM解码,提升推理准确性和计算效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15335 2026-01-23 cs.SE cs.AI cs.PL 85%

ToolCaching: Towards Efficient Caching for LLM Tool-calling

ToolCaching: 向LLM工具调用的高效缓存迈进

Yi Zhai, Dian Shen, Junzhou Luo, Bin Yang

机构 * School of Computer Science and Engineering, Southeast University(计算机科学与工程学院,东南大学)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 ToolCaching通过特征驱动和自适应的缓存框架,提升LLM工具调用的缓存命中率和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11833 2026-01-23 cs.AI 85%

Thought of Search: Planning with Language Models Through The Lens of Efficiency

思考搜索:通过效率的视角进行语言模型规划

Michael Katz, Harsha Kokel, Kavitha Srinivas, Shirin Sohrabi

机构 * David S. Hippocampus Department of Computer Science(戴维·S·海马科斯学院计算机科学系) Cranberry-Lemon University(Cranberry-Lemon大学) IBM Research(IBM研究院)

专题命中 效率与部署 :language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.AI

AI总结 本文提出了一种高效且保持正确性和完备性的LLM规划方法,通过解决四个代表性搜索问题展示其有效性,并呼吁研究社区关注效率与正确性的平衡。

Comments Accepted at NeurIPS 2024, https://papers.nips.cc/paper_files/paper/2024/hash/fa080fe0f218871faec1d8ba20e491d5-Abstract-Conference.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15528 2026-01-23 cs.DC cs.CR 85%

Securing LLM-as-a-Service for Small Businesses: An Industry Case Study of a Distributed Chatbot Deployment Platform

为中小企业保障LLM即服务的安全性:一个分布式聊天机器人部署平台的行业案例研究

Jiazhu Xie, Bowen Li, Heyu Fu, Chong Gao, Ziqi Xu, Fengling Han

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract)

AI总结 本文提出一个开源多租户平台,帮助中小企业低成本安全部署定制LLM聊天机器人,通过分布式集群和加密网络实现资源池化与隔离,同时集成防提示注入攻击机制。

Comments Accepted by AISC 2026

Journal ref Australasian Information Security Conference 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07526 2026-01-23 cs.SD cs.AI cs.CL cs.LG 85%

Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data

竞争性音频-语言模型:基于公共数据的高效单阶段训练

Gokul Karthik Kumar, Rishabh Saraf, Ludovick Lepauloux, Abdul Muneer, Billel Mokeddem, Hakim Hacid

专题命中 效率与部署 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 Falcon3-Audio基于少量公共数据实现高效单阶段训练,其1B模型在MMAU基准测试中表现优异,优于其他大型模型。

Comments Accepted at ASRU 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04037 2026-01-23 cs.AI 84%

Evolving in Tasks: Empowering the Multi-modality Large Language Model as the Computer Use Agent

任务演化:使多模态大语言模型成为计算机使用代理

Yuhao Cheng, Liang Tang, Shuxian Li, Yukang Huo, Tiaonan Duan, Kaer Huang, Yanzhe Jing, Yiqiang Yan

机构 * Lenovo Research(联想研究) China Agricultural University(中国农业大学)

专题命中 效率与部署 :large language model(title);language model(title);分类 cs.AI

AI总结 本文提出自演化代理(SEA),通过数据生成、强化学习和模型增强三个创新,使7B参数模型在计算机使用任务中超越同类模型并接近更大规模模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10903 2026-01-23 cs.NI cs.CL cs.LG 84%

LiLM-RDB-SFC: Lightweight Language Model with Relational Database-Guided DRL for Optimized SFC Provisioning

LiLM-RDB-SFC:轻量语言模型结合关系数据库引导的深度强化学习用于优化SFC配置

Parisa Fard Moshiri, Xinyu Zhu, Poonam Lohan, Burak Kantarci, Emil Janulewicz

机构 * University of Ottawa(渥太华大学) Ciena(CIENA公司)

专题命中 效率与部署 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.LG

AI总结 LiLM-RDB-SFC结合轻量语言模型与关系数据库引导深度强化学习,实现高效SFC配置,FLAN-T5在准确性与处理时间上优于BART和SQLCoder。

Comments 9 pages, 6 figures, Accepted to IEEE 16th International Conference on Network of the Future (NoF) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24592 2026-01-23 cs.AI cs.SE 83%

BPMN Assistant: An LLM-Based Approach to Business Process Modeling

BPMN助手:基于大语言模型的业务流程建模方法

Josip Tomo Licardo, Nikola Tankovic, Darko Etinger

机构 * Faculty of Informatics(信息学院) Juraj Dobrila University of Pula(朱拉·多布里拉大学)

专题命中 效率与部署 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 BPMN助手通过基于JSON的中间表示和大语言模型,提升BPMN图编辑效率与可靠性

Comments 22 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15710 2026-01-23 cs.AR cs.AI cs.LG 81%

FlexLLM: Composable HLS Library for Flexible Hybrid LLM Accelerator Design

FlexLLM:用于灵活混合LLM加速器设计的可组合HLS库

Jiahao Zhang, Zifan He, Nicholas Fraser, Michaela Blott, Yizhou Sun, Jason Cong

机构 * Computer Science, University of California, Los Angeles, California(计算机科学,加州大学洛杉矶分校)

专题命中 效率与部署 :LLM(title,abstract);分类 cs.AI、cs.LG

AI总结 FlexLLM通过可组合的HLS库实现灵活混合LLM加速器设计,提供高效量化和长上下文处理,显著提升性能与能效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15352 2026-01-23 cs.SE 80%

A Prompt-Based Framework for Loop Vulnerability Detection Using Local LLMs

基于提示的本地LLM循环漏洞检测框架

Adeyemi Adeseye, Aisvarya Adeseye

专题命中 效率与部署 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文提出基于本地LLM的循环漏洞检测框架,利用提示技术提高代码分析的准确性和安全性,验证Phi模型在性能上的优势。

Comments Accepted and Waiting to be published ICAI'25: 27th International Conference on Artificial Intelligence https://american-cse.org/csce2025/conferences-ICAI

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16011 2026-01-23 eess.IV cs.AI 79%

THOR: A Versatile Foundation Model for Earth Observation Climate and Society Applications

THOR:一种用于地球观测气候与社会应用的多功能基础模型

Theodor Forgaard, Jarle H. Reksten, Anders U. Waldeland, Valerio Marsocci, Nicolas Longépé, Michael Kampffmeyer, Arnt-Børre Salberg

机构 * Norwegian Computing Center(挪威计算中心) European Space Agency(欧洲航天局) UiT - The Arctic University of Tromsø(UiT-特罗姆斯大学) Φ \Phi -lab(Φ实验室)

专题命中 效率与部署 :foundation model(title,abstract);分类 cs.AI

AI总结 THOR是一种能够统一处理多种卫星数据的多功能基础模型,通过灵活的计算权衡提升在气候与社会应用中的性能。

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15504 2026-01-23 cs.LG q-bio.GN q-bio.QM 79%

SAGE-FM: A lightweight and interpretable spatial transcriptomics foundation model

SAGE-FM:一种轻量且可解释的空间转录组基础模型

Xianghao Zhan, Jingyu Xu, Yuanning Zheng, Zinaida Good, Olivier Gevaert

专题命中 效率与部署 :foundation model(title,abstract);分类 cs.LG

AI总结 SAGE-FM通过图卷积网络实现轻量级空间转录组学建模,能有效恢复被掩码基因并提升聚类和亚型预测性能。

Comments 26 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15578 2026-01-23 cs.NI cs.AI cs.LG cs.RO 79%

MapViT: A Two-Stage ViT-Based Framework for Real-Time Radio Quality Map Prediction in Dynamic Environments

MapViT:一种基于视觉Transformer的两阶段框架,用于动态环境中实时无线电质量地图预测

Cyril Shih-Huan Hsu, Xi Li, Lanfranco Zanzi, Zhiheng Yang, Chrysa Papagianni, Xavier Costa Pérez

机构 * Informatics Institute, University of Amsterdam(阿姆斯特丹大学信息学院) NEC Laboratories Europe(NEC欧洲实验室) i2CAT Foundation(i2CAT基金会) Catalan Institution for Research and Advanced Studies (ICREA)(加泰罗尼亚研究与高级研究机构(ICREA))

专题命中 效率与部署 :large language model(abstract);language model(abstract);foundation model(abstract);分类 cs.AI、cs.LG

AI总结 MapViT通过两阶段ViT框架实现实时动态环境中无线信号质量预测,结合预训练和微调方法提升准确性和效率,适用于资源受限的移动机器人场景。

Comments This paper has been accepted for publication at IEEE International Conference on Communications (ICC) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06193 2026-01-23 cs.CL cs.AI 79%

Do You Feel Comfortable? Detecting Hidden Conversational Escalation in AI Chatbots

你感觉舒适吗?检测AI聊天机器人中隐藏的对话升级

Jihyung Park, Saleh Afroogh, David Atkinson, Junfeng Jiao

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 效率与部署 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 GAUGE通过实时检测AI聊天机器人中隐藏的情感升级,提升对话中隐性伤害的识别能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11304 2026-01-23 cs.AI cs.CL cs.CV 79%

Leveraging Multimodal-LLMs Assisted by Instance Segmentation for Intelligent Traffic Monitoring

利用实例分割辅助的多模态大语言模型进行智能交通监控

Murat Arda Onsu, Poonam Lohan, Burak Kantarci, Aisha Syed, Matthew Andrews, Sean Kennedy

机构 * University of Ottawa(渥太华大学) Nokia Bell Labs(诺基亚贝尔实验室)

专题命中 效率与部署 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文利用多模态大语言模型和实例分割技术,实现高准确率的交通监控系统,提升交通管理效率和安全性。

Comments 6 pages, 7 figures, submitted to 30th IEEE International Symposium on Computers and Communications (ISCC) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18944 2026-01-23 cs.CV 78%

DINO in the Room: Leveraging 2D Foundation Models for 3D Segmentation

DINO在房间内:利用2D基础模型进行3D分割

Karim Knaebel, Kadir Yilmaz, Daan de Geus, Alexander Hermans, David Adrian, Timm Linder, Bastian Leibe

机构 * RWTH Aachen University(亚琛工业大学) Eindhoven University of Technology(埃因霍温理工大学) Bosch Center for AI(博世人工智能中心)

专题命中 效率与部署 :foundation model(title,abstract)

AI总结 DITR通过整合2D基础模型特征到3D分割模型中,实现了在室内外3D语义分割上的最佳性能。

Comments Accepted to 3DV 2026. Project page at https://vision.rwth-aachen.de/ditr

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15906 2026-01-23 cs.CV 78%

Opening the Black Box: Preliminary Insights into Affective Modeling in Multimodal Foundation Models

揭开黑箱:多模态基础模型中情感建模的初步洞察

Zhen Zhang, Runhao Zeng, Sicheng Zhao, Xiping Hu

机构 * Shenzhen MSU-BIT University, Shenzhen, China(深圳MSU-BIT大学) Tsinghua University, Beijing, China(清华大学)

专题命中 效率与部署 :foundation model(title,abstract)

AI总结 本研究揭示多模态基础模型中情感建模的关键机制,发现前馈门控投影模块对情感理解和生成至关重要,通过参数高效调整实现高性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15912 2026-01-23 cs.RO cs.AI 77%

TeNet: Text-to-Network for Compact Policy Synthesis

TeNet:基于文本的网络用于紧凑型策略合成

Ariyan Bighashdel, Kevin Sebastian Luck

机构 * Utrecht University(乌特勒支大学) Vrije Universiteit Amsterdam(阿姆斯特丹自由大学)

专题命中 效率与部署 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 TeNet通过文本条件化超网络实现紧凑型机器人策略合成,结合预训练语言模型的知识与高效执行,适用于实时资源受限的机器人控制任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15549 2026-01-23 cs.CV cs.AI 77%

VIOLA: Towards Video In-Context Learning with Minimal Annotations

VIOLA:迈向最小标注的视频情境学习

Ryo Fujii, Hideo Saito, Ryo Hachiuma

机构 * Keio University(Keio大学) Keio AI Research Center(Keio人工智能研究中心) NVIDIA(NVIDIA公司)

专题命中 效率与部署 :large language model(abstract);language model(abstract);prompting(abstract);分类 cs.AI

AI总结 VIOLA通过结合最小专家监督和大量未标注数据,实现高效视频情境学习,显著提升低资源环境下的适应性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24257 2026-01-23 cs.CR cs.LG 77%

VeriLLM: A Lightweight Framework for Publicly Verifiable Decentralized Inference

VeriLLM:一种轻量级的公开可验证去中心化推理框架

Ke Wang, Zishuo Zhao, Xinyuan Song, Zelin Li, Libin Xia, Chris Tong, Bill Shi, Wenjie Qu, Eric Yang, Lynn Ai

机构 * University of Illinois Urbana-Champaign, Gradient(伊利诺伊大学厄巴纳-香槟分校,Gradient) Emory University(埃默里大学) Ohio State University(俄亥俄州立大学) Peking University(北京大学) National University of Singapore(新加坡国立大学)

专题命中 效率与部署 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 VeriLLM通过轻量级经验重跑和链上检查实现公开可验证的去中心化LLM推理,提高效率与安全性,减少验证成本。

Comments 18 pages, 4 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16146 2026-01-23 cs.NI 75%

Low-altitude Multi-UAV-assisted Data Collection and Semantic Forwarding for Post-Disaster Relief

低空多无人机辅助的数据采集与语义转发用于灾后救援

Xiaoya Zheng, Geng Sun, Jiahui Li, Jiacheng Wang, Weijie Yuan, Qingqing Wu, Dusit Niyato, Abbas Jamalipour

专题命中 效率与部署 :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 本文提出了一种基于大语言模型的交替优化方法,用于低空多无人机辅助的数据采集与语义转发,以提高灾后救援中的通信效率和数据传输性能。

Comments 18 pages, 7 figures, journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10695 2026-01-23 cs.LG cs.AI cs.CL 75%

Introducing Verification Task of Set Consistency with Set-Consistency Energy Networks

引入集合一致性验证任务与集合一致性能量网络

Mooho Song, Hyeryung Son, Jay-Yoon Lee

机构 * Seoul National University(首尔国立大学)

专题命中 效率与部署 :LLM(abstract);prompting(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出集合一致性验证任务及SC-Energy模型,通过对比损失框架提升多陈述逻辑一致性验证性能,并发布新数据集

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025), Long Papers

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.05406 2026-01-23 cs.LG cs.CL 73%

Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes

人人都现在剪枝:仅前向传递的LLM结构剪枝

Steven Kolawole, Lucio Dery, Jean-François Kagy, Virginia Smith, Graham Neubig, Ameet Talwalkar

机构 * Carnegie Mellon University(卡内基梅隆大学) Google Research(谷歌研究)

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 Bonsai通过仅前向传递的扰动剪枝方法,实现了无需反向传播的高效结构化剪枝,显著降低内存和计算成本,同时提升模型压缩效率和性能。

Comments 19 pages, 6 fiigures, 16 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15676 2026-01-23 cs.SD cs.LG eess.AS 70%

Bridging the Perception Gap: A Lightweight Coarse-to-Fine Architecture for Edge Audio Systems

弥合感知差距:一种轻量级由粗到细架构用于边缘音频系统

Hengfan Zhang, Yueqian Lin, Hai Helen Li, Yiran Chen

机构 * Duke University(杜克大学)

专题命中 效率与部署 :LLM(abstract);language model(abstract);分类 cs.LG

AI总结 CoFi-Agent通过工具增强的条件边缘-云协作,在边缘音频系统中实现了更高效的感知与准确性提升。

Comments 10 pages, 3 figures, 2 tables. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15655 2026-01-23 cs.CV cs.AI 70%

Event-VStream: Event-Driven Real-Time Understanding for Long Video Streams

Event-VStream: 基于事件驱动的长视频流实时理解

Zhenghui Guo, Yuanbin Man, Junyuan Sheng, Bowen Lin, Ahmed Ahmed, Bo Jiang, Boyuan Zhang, Miao Yin, Sian Jin, Omprakash Gnawal, Chengming Zhang

机构 * University of Houston(德克萨斯大学休斯顿分校) The University of Texas at Arlington(德克萨斯理工大学) Indiana University Bloomington(印第安纳大学布卢明顿分校) Temple University(特拉华大学)

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 Event-VStream通过事件感知框架实现长视频流的实时理解,利用事件序列进行语义连贯的处理,提升性能并保持低延迟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15370 2026-01-23 cs.LG cs.AI 62%

Improving MoE Compute Efficiency by Composing Weight and Data Sparsity

通过组合权重和数据稀疏性提高MoE计算效率

Maciej Kilian, Oleg Mkrtchyan, Luke Zettlemoyer, Akshat Shrivastava, Armen Aghajanyan

机构 * University of Washington(华盛顿大学)

专题命中 效率与部署 :language model(abstract);分类 cs.AI、cs.LG

AI总结 通过结合权重和数据稀疏性,提高MoE计算效率,减少训练-推理不匹配,提升模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03592 2026-01-23 cs.CL cs.AI 62%

English K_Quantization of LLMs Does Not Disproportionately Diminish Multilingual Performance

LLM的英语量化并不 disproportionate 地损害多语言性能

Karl Audun Borgersen, Morten Goodwin

机构 * University of Agder(阿格德大学)

专题命中 效率与部署 :LLM(abstract);分类 cs.CL、cs.AI

AI总结 本文通过使用三种语言的重要性矩阵对Llama3.3 70B进行量化,发现当前量化实践不会不成比例地损害多语言性能。

Comments 8 pages, 6 figures, v2

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15809 2026-01-23 cs.CL 57%

SteerEval: Inference-time Interventions Strengthen Multilingual Generalization in Neural Summarization Metrics

SteerEval: 在推理时间干预增强神经摘要度量的多语言泛化

Silvia Casola, Ryan Soh-Eun Shim, Felicia Körner, Yuchen Mao, Barbara Plank

机构 * MaiNLP, Center for Information and Language Processing, LMU Munich(MaiNLP、信息与语言处理中心、慕尼黑大学) Language Science and Technology, Saarland University(语言科学与技术、萨尔兰大学)

专题命中 效率与部署 :language model(abstract);分类 cs.CL

AI总结 SteerEval通过在推理时干预激活向英语基准倾斜,提升多语言神经摘要度量的泛化能力。

Comments Submitted to ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15761 2026-01-23 cs.AI 57%

Off-Policy Actor-Critic with Sigmoid-Bounded Entropy for Real-World Robot Learning

基于Sigmoid受限熵的离策略Actor-Critic用于现实世界机器人学习

Xiefeng Wu, Mingyu Hu, Shu Zhang

机构 * Wuhan University(武汉大学)

专题命中 效率与部署 :pretraining(abstract);分类 cs.AI

AI总结 SigEnt-SAC通过Sigmoid受限熵机制,实现低成本、高效率的现实世界机器人强化学习。

Comments 7 pages main text 2 page reference

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11551 2026-01-23 cs.CV cs.IR cs.LG 57%

Multi-event Video-Text Retrieval

多事件视频-文本检索

Gengyuan Zhang, Jisen Ren, Jindong Gu, Volker Tresp

机构 * LMU Munich(慕尼黑大学) Munich Center for Machine Learning(慕尼黑机器学习中心) University of Oxford(牛津大学)

专题命中 效率与部署 :language model(abstract);分类 cs.LG

AI总结 本文提出多事件视频-文本检索任务,设计Me-Retriever模型,通过关键事件表示和新损失函数提升视频-文本检索性能。

Comments [fixed typos in equations] accepted to ICCV2023 Poster; some figures are not supported when viewed online, please download the file and view locally

详情

展开后加载摘要…

URL PDF HTML 收藏