arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-02-20 至 2026-02-20 共收录 157 信号源:cs.CL, cs.AI, cs.LG

1. 效率与部署 27 篇

2412.02039 2026-02-20 cs.CV cs.AI cs.LG 62%

Multi-View 3D Reconstruction using Knowledge Distillation

多视图3D重建使用知识蒸馏

Aditya Dutt, Ishikaa Lunawat, Manpreet Kaur

机构 * Stanford University(斯坦福大学)

专题命中 效率与部署 :foundation model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出使用知识蒸馏方法,通过Dust3r生成的3D重建点训练学生模型,以提升多视图3D重建的效率和性能,最终发现视觉Transformer架构在视觉和定量指标上表现最佳。

Comments 6 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17386 2026-02-20 cs.AI cs.IR 57%

Visual Model Checking: Graph-Based Inference of Visual Routines for Image Retrieval

视觉模型检查:基于图的图像检索中的视觉例行程序推断

Adrià Molina, Oriol Ramos Terrades, Josep Lladós

机构 * Centre de Visió per Computador(视觉计算中心) Universitat Autònoma de Barcelona(巴塞罗那自治大学)

专题命中 效率与部署 :pretraining(abstract);分类 cs.AI

AI总结 本文提出一种结合图基验证与神经代码生成的框架,用于提升图像检索中复杂查询的可信度和可验证性。

Comments Submitted for ICPR Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17102 2026-02-20 cs.LG 57%

Operationalization of Machine Learning with Serverless Architecture: An Industrial Operationalization of Machine Learning with Serverless Architecture: An Industrial Implementation for Harmonized System Code Prediction

基于无服务器架构的机器学习操作化:一种工业级的机器学习操作化:一种用于协调系统代码预测的工业实现

Sai Vineeth Kandappareddigari, Santhoshkumar Jagadish, Gauri Verma, Ilhuicamina Contreras, Christopher Dignam, Anmol Srivastava, Benjamin Demers

专题命中 效率与部署 :LLM(abstract);分类 cs.LG

AI总结 本文提出了一种基于无服务器架构的工业级机器学习操作化方案,用于协调系统代码预测,通过自动化流程提升准确性和成本效益。

Comments 13 pages. ICAD '26

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16738 2026-02-20 cs.MA cs.LG 57%

Self-Evolving Multi-Agent Network for Industrial IoT Predictive Maintenance

面向工业物联网预测性维护的自演化多智能体网络

Rebin Saleh, Khanh Pham Dinh, Balázs Villányi, Truong-Son Hy

机构 * Department of Electronics Technology, Faculty of Electrical Engineering and Informatics, Budapest University of Technology and Economics(布达佩斯技术与经济大学电子技术系) DataScienceWorld Department of Computer Science, The University of Alabama at Birmingham(阿拉巴马大学伯明翰分校计算机科学系)

专题命中 效率与部署 :LLM(abstract);分类 cs.LG

AI总结 SEMAS通过自演化多智能体架构实现工业物联网预测性维护的实时异常检测与资源优化

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 领域大模型 11 篇

2602.17418 2026-02-20 cs.AI 89%

A Privacy by Design Framework for Large Language Model-Based Applications for Children

为儿童使用的大语言模型应用设计隐私保护框架

Diana Addae, Diana Rogachova, Nafiseh Kahani, Masoud Barati, Michael Christensen, Chen Zhou

机构 * Computer Engineering Carleton University Ottawa, Canada(计算机工程学院卡罗伦大学加拿大) School of Information Technology Carleton University Ottawa, Canada(信息科技学院卡罗伦大学加拿大) Legal Studies Carleton University Ottawa, Canada(法律研究学院卡罗伦大学加拿大)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本文提出了一种基于隐私优先设计的框架,旨在为儿童使用的大语言模型应用提供隐私保护和法律合规性保障。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17410 2026-02-20 cs.IR cs.AI 89%

Improving LLM-based Recommendation with Self-Hard Negatives from Intermediate Layers

通过中间层自硬负样本改进基于LLM的推荐

Bingqian Li, Bowen Zheng, Xiaolei Wang, Long Zhang, Jinpeng Wang, Sheng Chen, Wayne Xin Zhao, Ji-rong Wen

机构 * GSAI, Renmin University of China(GSAI,中国人民大学)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);SFT(abstract)

AI总结 ILRec通过利用中间层自硬负样本改进LLM推荐系统,结合跨层优化和蒸馏提升负样本质量与判别性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16715 2026-02-20 cs.AI cs.CL cs.SY eess.SY 88%

Retrieval Augmented (Knowledge Graph), and Large Language Model-Driven Design Structure Matrix (DSM) Generation of Cyber-Physical Systems

检索增强(知识图谱),以及由大型语言模型驱动的面向网络系统的设计结构矩阵(DSM)生成

H. Sinan Bank, Daniel R. Herber

机构 * Department of Systems Engineering, Colorado State University(系统工程系,科罗拉多州立大学)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

AI总结 本文研究了利用大型语言模型、检索增强生成和图基RAG生成面向网络系统的设计结构矩阵,通过两个具体案例评估其在组件关系识别和生成方面的性能。

Comments 26 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16951 2026-02-20 eess.SP cs.LG 79%

BrainRVQ: A High-Fidelity EEG Foundation Model via Dual-Domain Residual Quantization and Hierarchical Autoregression

BrainRVQ: 一种通过双域残差量化和分层自回归的高保真EEG基础模型

Mingzhe Cui, Tao Chen, Yang Jiao, Yiqin Wang, Lei Xie, Yi Pan, Luca Mainardi

机构 * State Key Laboratory of Industrial Control Technology, Zhejiang University, Hangzhou, China(浙江大学工业控制技术状态重点实验室) Department of Electronics, Information and Bioengineering, Politecnico di Milano, Milan, Italy(米兰理工学院电子、信息与生物工程系) Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China(中国科学院深圳先进技术研究所)

专题命中 领域大模型 :foundation model(title,abstract);分类 cs.LG

AI总结 BrainRVQ通过双域残差量化和分层自回归方法,实现了对EEG信号的高保真基础模型,有效提升了细粒度神经表示的学习能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17308 2026-02-20 cs.AI cs.LG 79%

MedClarify: An information-seeking AI agent for medical diagnosis with case-specific follow-up questions

MedClarify:一种用于医学诊断的信息寻求AI代理,具有病例特定的后续问题

Hui Min Wong, Philip Heesen, Pascal Janetzky, Martin Bendszus, Stefan Feuerriegel

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 MedClarify通过生成病例特定的后续问题,提升医学诊断的准确性与可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17288 2026-02-20 cs.AI cs.CL 79%

ArXiv-to-Model: A Practical Study of Scientific LM Training

ArXiv-to-Model:科学语言模型训练的实践研究

Anuj Gupta

机构 * Independent Researcher(独立研究者)

专题命中 领域大模型 :large language model(abstract);language model(abstract);pretraining(abstract);分类 cs.CL、cs.AI

AI总结 本研究通过从arXiv LaTeX源数据直接训练科学语言模型,探讨了预处理、分词和计算资源对模型训练的影响,提供了工程层面的训练实践与透明分析。

Comments 15 pages, 6 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17450 2026-02-20 cs.IR cs.AI 70%

Beyond Pipelines: A Fundamental Study on the Rise of Generative-Retrieval Architectures in Web Research

超越流水线:生成-检索架构在网页研究中的崛起根本研究

Amirereza Abbasi, Mohsen Hooshmand

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文研究了生成-检索架构在网页研究中的崛起,探讨了LLMs通过RAG对网页研究和行业的影响,分析了关键进展、挑战及未来方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15943 2026-02-20 cs.CV 67%

Boosting Medical Visual Understanding From Multi-Granular Language Learning

通过多粒度语言学习提升医学视觉理解

Zihan Li, Yiqing Wang, Sina Farsiu, Paul Kinahan

机构 * University of Washington(华盛顿大学) Duke University(杜克大学)

专题命中 领域大模型 :language model(abstract);pretraining(abstract)

AI总结 本文提出MGLL框架,通过多粒度语言学习提升医学影像的视觉理解能力,改进多标签和跨粒度对齐,提升下游任务性能。

Comments Accepted by ICLR 2026. 40 pages

Journal ref The Fourteenth International Conference on Learning Representations (ICLR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17607 2026-02-20 cs.AI cs.LG cs.NA math.NA 62%

AutoNumerics: An Autonomous, PDE-Agnostic Multi-Agent Pipeline for Scientific Computing

AutoNumerics: 一种自主的、与偏微分方程无关的多智能体流水线用于科学计算

Jianda Du, Youran Sun, Haizhao Yang

机构 * University of Maryland(马里兰大学)

专题命中 领域大模型 :LLM(abstract);分类 cs.AI、cs.LG

AI总结 AutoNumerics通过多智能体框架自主设计并验证数值求解器,基于自然语言描述实现高效准确的偏微分方程求解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17535 2026-02-20 cs.CV 50%

LATA: Laplacian-Assisted Transductive Adaptation for Conformal Uncertainty in Medical VLMs

LATA:拉普拉斯辅助的转导适应用于医学视觉语言模型中的符合不确定性

Behzad Bozorgtabar, Dwarikanath Mahapatra, Sudipta Roy, Muzammal Naseer, Imran Razzak, Zongyuan Ge

机构 * Aarhus University ( A3 Lab )(奥胡斯大学) Khalifa University(哈利法大学) Jio Institute(乔研究所) MBZUAI(穆萨大学人工智能研究所) Monash University(墨尔本大学)

专题命中 领域大模型 :language model(abstract)

AI总结 LATA通过拉普拉斯辅助的转导适应方法,提升医学VLM在域转移下的不确定性校准,减少预测集规模和类间覆盖不平衡,同时保持高覆盖效率。

Comments 18 pages, 6 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17215 2026-02-20 cs.HC 50%

NotebookRAG: Retrieving Multiple Notebooks to Augment the Generation of EDA Notebooks for Crowd-Wisdom

NotebookRAG: 通过检索多个笔记本增强 crowdsourcing 智慧的 EDA 笔记本生成

Yi Shan, Yixuan He, Zekai Shao, Kai Xu, Siming Chen

专题命中 领域大模型 :LLM(abstract)

AI总结 NotebookRAG通过检索多个笔记本增强crowdsourcing智慧,以生成高质量的EDA笔记本。

Comments 11 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 知识编辑与模型理解 7 篇

2602.17529 2026-02-20 cs.AI 89%

Enhancing Large Language Models (LLMs) for Telecom using Dynamic Knowledge Graphs and Explainable Retrieval-Augmented Generation

通过动态知识图谱和可解释检索增强生成增强大型语言模型用于电信

Dun Yuan, Hao Zhou, Xue Liu, Hao Chen, Yan Xin, Jianzhong, Zhang

机构 * School of Computer Science, McGill University(麦吉尔大学计算机科学学院) Standards and Mobility Innovation Lab, Samsung Research America(三星美国标准与移动创新实验室)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本文提出KG-RAG框架,通过整合知识图谱与检索增强生成技术,提升大型语言模型在电信领域中的准确性与可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16977 2026-02-20 cs.LG cs.CR 89%

Fail-Closed Alignment for Large Language Models

大型语言模型的Fail-Closed对齐

Zachary Coalson, Beth Sohler, Aiden Gabriel, Sanghyun Hong

机构 * Oregon State University, Corvallis, OR USA(俄勒冈州立大学)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.LG

AI总结 本文提出fail-closed对齐原则,通过冗余独立路径提升LLM安全,有效对抗基于提示的Jailbreak攻击。

Comments Pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16740 2026-02-20 cs.LG cs.AI 81%

Quantifying LLM Attention-Head Stability: Implications for Circuit Universality

量化大语言模型注意力头的稳定性:对电路通用性的启示

Karan Bali, Jack Stanley, Praneet Suresh, Danilo Bzdok

机构 * Mila - Quebec Artificial Intelligence Institute(魁北克人工智能研究所) McGill University(麦吉尔大学)

专题命中 知识编辑与模型理解 :LLM(title);language model(abstract);分类 cs.AI、cs.LG

AI总结 本研究量化了大语言模型注意力头的稳定性,发现中层头最不稳定但表示最独特,深层模型中层分歧更明显,权重衰减优化提升稳定性,残差流较稳定,为AI系统白盒监控提供新视角。

Comments Main Body: 8 pages, Total length: 33 pages, Code repo: https://github.com/karanbali/attention_head_seed_stability , Weights repo: https://huggingface.co/karanbali/attention_head_seed_stability

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17532 2026-02-20 q-bio.GN cs.AI 79%

Systematic Evaluation of Single-Cell Foundation Model Interpretability Reveals Attention Captures Co-Expression Rather Than Unique Regulatory Signal

单细胞基础模型解释性系统评估揭示注意力捕捉共表达而非独特调控信号

Ihor Kendiukhov

机构 * Department of Computer Science University of Tübingen(计算机科学系图宾根大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.AI

AI总结 单细胞基础模型中注意力机制捕捉共表达而非独特调控信号,CSSI提升GRN恢复效率

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16980 2026-02-20 cs.LG cs.CR 79%

Discovering Universal Activation Directions for PII Leakage in Language Models

发现语言模型中PII泄露的通用激活方向

Leo Marchyok, Zachary Coalson, Sungho Keum, Sooel Son, Sanghyun Hong

机构 * Oregon State University, Corvallis OR, USA Korea Advanced Institute of Science \& Technology, Daejeon, South Korea

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.LG

AI总结 UniLeak通过识别语言模型中通用激活方向,提升PII泄露可能性,为隐私安全研究提供新视角。

Comments Pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17256 2026-02-20 physics.soc-ph physics.app-ph physics.comp-ph 75%

On the Concept of Violence: A Comparative Study of Human and AI Judgments

关于暴力概念的观念:人类与AI判断的比较研究

Mariachiara Stellato, Francesco Lancia, Chiara Galeazzi, Nico Curti

专题命中 知识编辑与模型理解 :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 本文通过比较人类与AI对暴力的判断,探讨了AI如何处理模糊的道德概念及人类解释的转化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17259 2026-02-20 cs.RO 67%

FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment

FRAPPE:通过多未来表示对齐将世界建模注入通用策略

Han Zhao, Jingbo Wang, Wenxuan Song, Shuai Chen, Yang Liu, Yan Wang, Haoang Li, Donglin Wang

机构 * Zhejiang University(浙江大学) Westlake University(西湖大学) South China University of Technology(华南理工大学) ShanghaiTech University(上海科技大学) Tsinghua University(清华大学)

专题命中 知识编辑与模型理解 :foundation model(abstract);post-training(abstract)

AI总结 FRAPPE通过多未来表示对齐提升通用机器人策略的世界建模能力,提高微调效率并减少数据依赖,实验证明其在长时序和未见场景中的泛化性能优于现有方法。

Comments Project Website: https://h-zhao1997.github.io/frappe

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他LLM 15 篇

2511.18696 2026-02-20 cs.CL cs.AI 92%

Empathetic Cascading Networks: A Multi-Stage Prompting Technique for Reducing Social Biases in Large Language Models

共情级联网络:一种减少大型语言模型社会偏见的多阶段提示技术

Wangjiaxuan Xin

专题命中 其他LLM :large language model(title,abstract);language model(title,abstract);prompting(title,abstract);分类 cs.CL、cs.AI

AI总结 ECN通过多阶段提示技术减少大型语言模型的社会偏见,提升对话AI的共情与包容性。

Comments Further revision on experiments and pipeline design

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13469 2026-02-20 cs.HC cs.AI 88%

How Multimodal Large Language Models Support Access to Visual Information: A Diary Study With Blind and Low Vision People

多模态大语言模型如何支持视障人士获取视觉信息:一项与盲人和低视力人士的日记研究

Ricardo E. Gonzalez Penuela, Crescentia Jung, Sharon Y Lin, Ruiying Hu, Shiri Azenkot

机构 * Cornell University(康奈尔大学)

专题命中 其他LLM :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本研究探讨了多模态大语言模型如何通过视觉助手技能支持视障人士获取视觉信息,并发现其在实际应用中的表现及改进方向。

Comments 24 pages, 17 figures, 7 tables, appendix section, to appear main track CHI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17481 2026-02-20 cs.HC 85%

ShadAR: LLM-driven shader generation to transform visual perception in Augmented Reality

ShadAR: 基于大语言模型的着色器生成以在增强现实中转变视觉感知

Yanni Mei, Samuel Wendt, Florian Mueller, Jan Gugenheimer

专题命中 其他LLM :LLM(title,abstract);large language model(abstract);language model(abstract)

AI总结 ShadAR利用大语言模型生成着色器,实现增强现实中的视觉感知实时转变,提升包容性和创造性。

Journal ref 2025 IEEE International Symposium on Mixed and Augmented Reality Adjunct (ISMAR-Adjunct), Daejeon, Korea, Republic of, 2025, pp. 959-960

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17483 2026-02-20 cs.HC cs.AI cs.CL cs.CY 79%

What Do LLMs Associate with Your Name? A Human-Centered Black-Box Audit of Personal Data

LLMs与你的名字关联什么?对个人数据的人本黑盒审计

Dimitri Staufer, Kirsten Morehouse

机构 * Columbia University(哥伦比亚大学)

专题命中 其他LLM :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 研究通过LMP2工具审计LLMs对个人数据的关联,发现模型能准确生成用户特征,引发对数据隐私权在LLMs中的适用性讨论。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17293 2026-02-20 cs.CY 75%

Human attribution of empathic behaviour to AI systems

人类将共情行为归因于人工智能系统

Jonas Festor, Ivo Snels, Bennett Kleinberg

专题命中 其他LLM :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 研究探讨了人类和AI生成建议在共情感知上的差异,发现LLM生成内容在共情方面表现更优,但作者标签对情感共情影响有限。

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09201 2026-02-20 cs.LG cs.AI cs.CL 75%

Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs

多模态提示优化:为何不利用多种模态为大语言模型服务

Yumin Choi, Dongki Kim, Jinheon Baek, Sung Ju Hwang

机构 * KAIST(韩国科学技术院)

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出多模态提示优化问题,提出MPO框架,通过联合优化和贝叶斯策略提升多模态提示效果,验证其在图像、视频等多模态任务中的优越性。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16997 2026-02-20 cs.SE cs.AI cs.CL 73%

Exploring LLMs for User Story Extraction from Mockups

探索LLMs用于从mockups中提取用户故事

Diego Firmenich, Leandro Antonelli, Bruno Pazos, Fabricio Lozada, Leonardo Morales

机构 * Departamento de Informática, Facultad de Ingeniería, Universidad Nacional de la Patagonia, Argentina(阿根廷国立巴塔哥尼亚大学计算机系) LIFIA, Facultad de Informática, Universidad Nacional de La Plata, Argentina(阿根廷国立拉普拉塔大学计算机学院) CAETI, Facultad de Tecnología Informática - Universidad Abierta Interamericana(拉丁美洲开放大学信息技术学院) Laboratorio de Ciencias de la Imágenes, DIEC, UNS(图像科学实验室) Instituto Patagónico de Ciencias Sociales y Humanas, CONICET-CENPAT(巴塔哥尼亚社会科学与人类科学研究所) UNIANDES, Universidad Regional Autónoma de los Andes, Ecuador(厄瓜多尔安第斯地区自治大学)

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文探讨了利用LLMs结合LEL从mockups中自动提取用户故事,提升了生成准确性与适用性,推动AI在需求工程中的应用。

Comments 14 pages, 6 figures. Preprint of the paper published in the 28th Workshop on Requirements Engineering (WER 2025)

Journal ref Proceedings of the 28th Workshop on Requirements Engineering (WER2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17645 2026-02-20 cs.LG cs.AI cs.CL cs.CV 67%

Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting

通过细粒度细节靶向推动大视觉-语言模型攻击的前沿

Xiaohan Zhao, Zhaoyi Li, Yaxin Luo, Jiacheng Cui, Zhiqiang Shen

专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 通过细粒度细节靶向改进M-Attack,显著提升黑盒攻击效果,成功率达100%

Comments Code at: https://github.com/vila-lab/M-Attack-V2

详情

展开后加载摘要…

URL PDF HTML 收藏