arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 12429 信号源:cs.CL, cs.AI, cs.LG

1. 预训练与数据 12429 篇

2509.02602 2025-09-04 eess.IV 78%

Masked Autoencoder Pretraining and BiXLSTM ResNet Architecture for PET/CT Tumor Segmentation

Moona Mazher, Steven A Niederer, Abdul Qayyum

专题命中 预训练与数据 :pretraining(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15798 2025-08-26 cs.CV 78%

MM-Retinal V2: Transfer an Elite Knowledge Spark into Fundus Vision-Language Pretraining

Ruiqi Wu, Na Su, Chenran Zhang, Tengfei Ma, Tao Zhou, Zhiting Cui, Nianfeng Tang, Tianyu Mao, Yi Zhou, Wen Fan, Tianxing Wu, Shenqi Jing, Huazhu Fu

机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Department of Ophthalmology, The First Affiliated Hospital of Nanjing Medical University(南京医科大学第一附属医院眼科部) School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院) Institute of High-Performance Computing, Agency for Science, Technology and Research(科技研究局高性能计算研究所)

专题命中 预训练与数据 :pretraining(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.02943 2025-08-26 cs.CR cs.AI cs.CL cs.LG 78%

PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding

Krishna Kanth Nakka, Ahmed Frikha, Ricardo Mendes, Xue Jiang, Xuebing Zhou

机构 * Huawei Munich Research Center(华为慕尼黑研究中心)

专题命中 预训练与数据 :LLM(title);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at PrivateNLP Workshop at ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07144 2025-08-12 cs.CV 78%

Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models

Xuanhan Wang, Huimin Deng, Ke Liu, Jun Wang, Lianli Gao, Jingkuan Song

机构 * Tongji University(同济大学) University of Electronic Science and Technology of China(电子科技大学)

专题命中 预训练与数据 :pretraining(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14683 2025-07-29 cs.CV 78%

Emerging Properties in Unified Multimodal Pretraining

Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou, Feng Li, Zeyu Wang, Shu Zhong, Weihao Yu, Xiaonan Nie, Ziang Song, Guang Shi, Haoqi Fan

机构 * ByteDance Seed(字节跳动种子) Shenzhen Institutes of Advanced Technology(深圳先进技术研究院) Monash University(墨尔本大学) Hong Kong University of Science and Technology(香港科学与技术大学) UC Santa Cruz(加州大学圣克ruz分校)

专题命中 预训练与数据 :pretraining(title,abstract)

Comments 37 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06132 2025-07-23 cs.CV 78%

USP: Unified Self-Supervised Pretraining for Image Generation and Understanding

Xiangxiang Chu, Renda Li, Yong Wang

机构 * AMAP, Alibaba Group(阿里集团)

专题命中 预训练与数据 :pretraining(title,abstract)

Comments Accepted to ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15285 2025-07-22 cs.CV 78%

In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems

Lazaro Janier Gonzalez-Soler, Maciej Salwowski, Christoph Busch

机构 * da/sec - Biometrics and Security Research Group(da/sec生物识别与安全研究组) Technical University of Denmark(丹麦技术大学)

专题命中 预训练与数据 :language model(title,abstract)

Comments Submitted to IEEE-TIFS

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07886 2025-07-22 cs.CV 78%

EgoM2P: Egocentric Multimodal Multitask Pretraining

Gen Li, Yutong Chen, Yiqian Wu, Kaifeng Zhao, Marc Pollefeys, Siyu Tang

机构 * ETH Zürich(苏黎世联邦理工学院) Zhejiang University(浙江大学) Microsoft(微软公司)

专题命中 预训练与数据 :pretraining(title);foundation model(abstract)

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10306 2025-07-15 cs.CV 78%

Contrastive Pretraining with Dual Visual Encoders for Gloss-Free Sign Language Translation

Ozge Mercanoglu Sincan, Richard Bowden

机构 * CVSSP, University of Surrey(CVSSP,英国萨里大学)

专题命中 预训练与数据 :pretraining(title,abstract)

Comments Accepted at 9th Workshop on Sign Language Translation and Avatar Technologies (SLTAT), will be held in conjunction with IVA'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09537 2025-07-15 cs.RO 78%

Self-supervised Pretraining for Integrated Prediction and Planning of Automated Vehicles

Yangang Ren, Guojian Zhan, Chen Lv, Jun Li, Fenghua Liang, Keqiang Li

机构 * Changan Automobile(长安汽车) Tsinghua University(清华大学) Nanyang Technological University(南洋理工大学)

专题命中 预训练与数据 :pretraining(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07802 2025-07-14 cs.CV 78%

Synergistic Prompting for Robust Visual Recognition with Missing Modalities

Zhihui Zhang, Luanyuan Dai, Qika Lin, Yunfeng Diao, Guangyin Jin, Yufei Guo, Jing Zhang, Xiaoshuai Hao

机构 * Beijing Institute of Technology(北京理工大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Nanjing University of Science and Technology(南京理工大学) National University of Singapore(新加坡国立大学) Systems Laboratory of Anhui Province, Hefei University of Technology(安徽省系统实验室,合肥工业大学) National Innovative Institute of Defense Technology(国家创新防御技术研究院) Intelligent Science & Technology Academy of CASIC(中国航天科技集团智能科学与技术学院) School of Computer Science, Wuhan University(武汉大学计算机学院)

专题命中 预训练与数据 :prompting(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06761 2025-07-10 cs.CV 78%

Finetuning Vision-Language Models as OCR Systems for Low-Resource Languages: A Case Study of Manchu

Yan Hon Michael Chung, Donghyeok Choi

机构 * Division of Humanities(人文学院) The Hong Kong University of Science and Technology(香港科学与技术大学) History, Religious and Philosophy Academy(历史、宗教与哲学学院) Hong Kong Baptist University(香港 Baptist 大学)

专题命中 预训练与数据 :language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18709 2025-07-10 cs.CV 78%

Revisiting Automatic Data Curation for Vision Foundation Models in Digital Pathology

Boqi Chen, Cédric Vincent-Cuaz, Lydia A. Schoenpflug, Manuel Madeira, Lisa Fournier, Vaishnavi Subramanian, Sonali Andani, Samuel Ruiperez-Campillo, Julia E. Vogt, Raphaëlle Luisier, Dorina Thanou, Viktor H. Koelzer, Pascal Frossard, Gabriele Campanella, Gunnar Rätsch

机构 * Dept. of Computer Science, ETH Zurich, Zurich, Switzerland(苏黎世联邦理工学院计算机科学系) AI Center, ETH Zurich, Zurich, Switzerland(苏黎世联邦理工学院人工智能中心) Signal Processing Laboratory (LTS4), EPFL, Lausanne, Switzerland(日内瓦联邦理工学院信号处理实验室) University of Basel, Basel, Switzerland(巴塞尔大学) University Hospital of Basel, Basel, Switzerland(巴塞尔大学医院) Idiap Research Institute, Martigny, Switzerland(伊迪普研究 institute) Icahn School of Medicine at Mount Sinai, New York, United States(辛辛那提医学中心伊坎医学院)

专题命中 预训练与数据 :foundation model(title,abstract)

Comments MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00676 2025-07-02 cs.CV 78%

A Unified Transformer-Based Framework with Pretraining For Whole Body Grasping Motion Generation

Edward Effendy, Kuan-Wei Tseng, Rei Kawakami

机构 * Institute of Science Tokyo(东京科学研究院)

专题命中 预训练与数据 :pretraining(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23259 2025-07-01 eess.IV cs.CV 78%

Improving Myocardial Infarction Detection via Synthetic ECG Pretraining

Lachin Naghashyar

机构 * Department of Computer Science, University of Oxford, Oxford, UK(计算机科学系,牛津大学,牛津,英国)

专题命中 预训练与数据 :pretraining(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19590 2025-06-25 eess.IV cs.CV 78%

Learning from Anatomy: Supervised Anatomical Pretraining (SAP) for Improved Metastatic Bone Disease Segmentation in Whole-Body MRI

Joris Wuts, Jakub Ceranka, Nicolas Michoux, Frédéric Lecouvet, Jef Vandemeulebroucke

机构 * Vrije Universiteit Brussel(瓦布伦大学布鲁塞尔分校) UCLouvain(列日大学) imec Cliniques universitaires Saint Luc & Institut de Recherche Expérimentale et Clinique (IREC)(圣卢大学医院及实验与临床研究所) Universitair Ziekenhuis Brussel(布鲁塞尔大学医院)

专题命中 预训练与数据 :pretraining(title,abstract)

Comments This preprint is currently under review at *Computers in Biology and Medicine* (Elsevier). This version has not been peer-reviewed

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10538 2025-06-25 physics.chem-ph 78%

Foundation Models for Atomistic Simulation of Chemistry and Materials

Eric C. -Y. Yuan, Yunsheng Liu, Junmin Chen, Peichen Zhong, Sanjeev Raja, Tobias Kreiman, Santiago Vargas, Wenbin Xu, Martin Head-Gordon, Chao Yang, Samuel M. Blau, Bingqing Cheng, Aditi Krishnapriyan, Teresa Head-Gordon

专题命中 预训练与数据 :foundation model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17815 2025-06-24 cs.SD eess.AS 78%

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding

Julien Guinot, Alain Riou, Elio Quinton, György Fazekas

专题命中 预训练与数据 :pretraining(title,abstract)

Comments Accepted to ISMIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09042 2025-06-19 cs.CV 78%

Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models

Xuanchi Ren, Yifan Lu, Tianshi Cao, Ruiyuan Gao, Shengyu Huang, Amirmojtaba Sabour, Tianchang Shen, Tobias Pfaff, Jay Zhangjie Wu, Runjian Chen, Seung Wook Kim, Jun Gao, Laura Leal-Taixe, Mike Chen, Sanja Fidler, Huan Ling

专题命中 预训练与数据 :foundation model(title,abstract)

Comments Only the core contributors are listed. The full list of contributors can be found in Appendix A of this paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11171 2025-06-19 cs.CL cs.AI cs.LG 78%

LLäMmlein: Transparent, Compact and Competitive German-Only Language Models from Scratch

Jan Pfister, Julia Wunderle, Andreas Hotho

机构 * Data Science Chair Center for Artificial Intelligence and Data Science (CAIDAS)(数据科学主席中心人工智能与数据科学中心(CAIDAS)) Julius-Maximilians-Universität Würzburg (JMU)(乌尔姆-马克斯-普朗克大学(JMU))

专题命中 预训练与数据 :language model(title);分类 cs.CL、cs.AI、cs.LG

Comments camera ready @ACL25; https://www.informatik.uni-wuerzburg.de/datascience/projects/nlp/llammlein/

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15289 2025-06-18 cs.CV 78%

Learning Invariant Causal Mechanism from Vision-Language Models

Zeen Song, Siyu Zhao, Xingyu Zhang, Jiangmeng Li, Changwen Zheng, Wenwen Qiang

机构 * Institute of Software Chinese Academy of Sciences, Beijing, China(中国科学院软件研究所) University of the Chinese Academy of Sciences(中国科学院大学)

专题命中 预训练与数据 :language model(title);pretraining(abstract)

Comments Accepted to ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13414 2025-06-17 eess.AS 78%

BUT System for the MLC-SLM Challenge

Alexander Polok, Jiangyu Han, Dominik Klement, Samuele Cornell, Jan Černocký, Lukáš Burget

专题命中 预训练与数据 :SLM(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17549 2025-06-17 cs.RO 78%

Canonical Representation and Force-Based Pretraining of 3D Tactile for Dexterous Visuo-Tactile Policy Learning

Tianhao Wu, Jinzhou Li, Jiyao Zhang, Mingdong Wu, Hao Dong

机构 * Center on Frontiers of Computing Studies, School of Computer Science, Peking University(前沿计算研究中心,计算机科学学院,北京大学) PKU-Agibot Lab, School of Computer Science, Peking University(北京大学计算机科学学院PKU-Agibot实验室) National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(国家多媒体信息处理重点实验室,计算机科学学院,北京大学)

专题命中 预训练与数据 :pretraining(title,abstract)

Comments Accepted to ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06281 2025-06-09 cs.CV 78%

TerraFM: A Scalable Foundation Model for Unified Multisensor Earth Observation

Muhammad Sohail Danish, Muhammad Akhtar Munir, Syed Roshaan Ali Shah, Muhammad Haris Khan, Rao Muhammad Anwer, Jorma Laaksonen, Fahad Shahbaz Khan, Salman Khan

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学) University College London(伦敦大学学院) Aalto University(艾尔沃斯大学) Linköping University, Sweden(瑞典林奈大学) Australian National University(澳大利亚国立大学)

专题命中 预训练与数据 :foundation model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01724 2025-06-03 cs.CV 78%

Active Learning via Vision-Language Model Adaptation with Open Data

Tong Wang, Jiaqi Wang, Shu Kong

机构 * University of Macau(澳门大学) Shanghai AI Lab(上海人工智能实验室) Institute of Collaborative Innovation project webpage(协同创新项目研究所)

专题命中 预训练与数据 :language model(title);pretraining(abstract)

Comments Here is the project webpage: https://leowangtong.github.io/ALOR/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24545 2025-06-02 eess.AS cs.SD 78%

Pretraining Multi-Speaker Identification for Neural Speaker Diarization

Shota Horiguchi, Atsushi Ando, Marc Delcroix, Naohiro Tawara

专题命中 预训练与数据 :pretraining(title,abstract)

Comments Accepted to Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22072 2025-05-29 cs.SD eess.AS 78%

On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition

Shujie HU, Xurong Xie, Mengzhe Geng, Jiajun Deng, Huimeng Wang, Guinan Li, Chengxi Deng, Tianzi Wang, Mingyu Cui, Helen Meng, Xunying Liu

机构 * The Chinese University of Hong Kong(香港中文大学) Chinese Academy of Sciences(中国科学院) China National Research Council Canada(中国国家研究委员会加拿大) Institute of Software(软件研究所)

专题命中 预训练与数据 :foundation model(title,abstract)

Comments Accepted by Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.06510 2025-05-27 cs.CV 78%

Can Large Vision-Language Models Correct Semantic Grounding Errors By Themselves?

Yuan-Hong Liao, Rafid Mahmood, Sanja Fidler, David Acuna

机构 * University of Toronto(多伦多大学) Vector Institute(向量研究所) NVIDIA(英伟达) University of Ottawa(渥太华大学)

专题命中 预训练与数据 :language model(title,abstract)

Comments Accepted at CVPR 2025. 22 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19437 2025-05-27 cs.SD eess.AS 78%

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval

Haoqin Sun, Jingguang Tian, Jiaming Zhou, Hui Wang, Jiabei He, Shiwan Zhao, Xiangyu Kong, Desheng Hu, Xinkang Xu, Xinhui Hu, Yong Qin

机构 * Nankai University(南开大学) TMCC, College of Computer Science(TMCC计算机学院) Hithink RoyalFlush AI Research Institute(Hithink RoyalFlush人工智能研究院) University of Exeter(埃克塞特大学)

专题命中 预训练与数据 :pretraining(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08956 2025-05-20 cs.CR cs.SE 78%

Defending Code Language Models against Backdoor Attacks with Deceptive Cross-Entropy Loss

Guang Yang, Yu Zhou, Xiang Chen, Xiangyu Zhang, Terry Yue Zhuo, David Lo, Taolue Chen

专题命中 预训练与数据 :language model(title,abstract)

Comments TOSEM

详情

展开后加载摘要…

URL PDF HTML 收藏