arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 140189 信号源:cs.CL, cs.AI, cs.LG

1. 预训练与数据 12486 篇

2510.13913 2025-10-17 cs.CL cs.AI 62%

Synthesizing Agentic Data for Web Agents with Progressive Difficulty Enhancement Mechanisms

Shrey Pandit, Xuan-Phi Nguyen, Yifei Ming, Austin Xu, Jiayu Wang, Caiming Xiong, Shafiq Joty

机构 * Salesforce AI Research(Salesforce AI研究院) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 预训练与数据 :language model(abstract);分类 cs.CL、cs.AI

Comments Preprint. ICLR 26 submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03052 2025-10-16 cs.CL cs.LG 62%

Teaching Models to Understand (but not Generate) High-risk Data

Ryan Wang, Matthew Finlayson, Luca Soldaini, Swabha Swayamdipta, Robin Jia

机构 * Department of Computer Science, University of Southern California(计算机科学系,南加州大学) Allen Institute for AI(人工智能研究所)

专题命中 预训练与数据 :language model(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18966 2025-10-15 cs.LG cs.AI q-bio.BM 62%

Protein Design with Dynamic Protein Vocabulary

Nuowei Liu, Jiahao Kuang, Yanting Liu, Tao Ji, Changzhi Sun, Man Lan, Yuanbin Wu

机构 * School of Computer Science and Technology, East China Normal University(东华师范大学计算机科学与技术学院) College of Foreign Languages and Literatures, Fudan University(复旦大学外国语言文学学院) Institute of Artificial Intelligence (TeleAI), China Telecom(中国电信人工智能研究所)

专题命中 预训练与数据 :language model(abstract);分类 cs.AI、cs.LG

Comments Accepted to NeurIPS 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09687 2025-10-14 cs.LG cs.AI 62%

On the Occurence of Critical Learning Periods in Neural Networks

Stanisław Pawlak

机构 * Stanisław Pawlak(斯坦尼斯劳·帕瓦克)

专题命中 预训练与数据 :pretraining(abstract);分类 cs.AI、cs.LG

Comments 8 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09596 2025-10-13 cs.LG cs.AI 62%

BaNEL: Exploration Posteriors for Generative Modeling Using Only Negative Rewards

Sangyun Lee, Brandon Amos, Giulia Fanti

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 预训练与数据 :post-training(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00885 2025-10-10 cs.CL cs.LG 62%

Scaling Laws Are Unreliable for Downstream Tasks: A Reality Check

Nicholas Lourie, Michael Y. Hu, Kyunghyun Cho

机构 * New York University(纽约大学) Prescient Design CIFAR LMB

专题命中 预训练与数据 :pretraining(abstract);分类 cs.CL、cs.LG

Comments EMNLP Findings 2025 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08113 2025-10-10 cs.LG cs.AI cs.MA 62%

Bayesian Decision Making around Experts

Daniel Jarne Ornia, Joel Dyer, Nicholas Bishop, Anisoara Calinescu, Michael Wooldridge

机构 * University of Oxford(牛津大学)

专题命中 预训练与数据 :pretraining(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06949 2025-10-09 cs.LG cs.AI 62%

Grouped Differential Attention

Junghwan Lim, Sungmin Lee, Dongseok Kim, Wai Ting Cheung, Beomgyu Kim, Taehwan Kim, Haesol Lee, Junhyeok Lee, Dongpin Oh, Eunhwan Park

机构 * Motif Technologies(Motif科技)

专题命中 预训练与数据 :pretraining(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19823 2025-10-08 cs.LG cs.AI 62%

Persona Features Control Emergent Misalignment

Miles Wang, Tom Dupré la Tour, Olivia Watkins, Alex Makelov, Ryan A. Chi, Samuel Miserendino, Jeffrey Wang, Achyuta Rajaram, Johannes Heidecke, Tejal Patwardhan, Dan Mossing

机构 * OpenAI

专题命中 预训练与数据 :language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11004 2025-10-08 cs.CL cs.AI 62%

Illusion or Algorithm? Investigating Memorization, Emergence, and Symbolic Processing in In-Context Learning

Jingcheng Niu, Subhabrata Dutta, Ahmed Elshabrawy, Harish Tayyar Madabushi, Iryna Gurevych

机构 * Technical University of Darmstadt(德累斯顿技术大学) UKP Lab(UKP实验室) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) The University of Bath(巴斯大学)

专题命中 预训练与数据 :language model(abstract);分类 cs.CL、cs.AI

Comments TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04945 2025-10-07 cs.CL cs.AI 62%

A First Context-Free Grammar Applied to Nawatl Corpora Augmentation

Juan-José Guzmán-Landa, Juan-Manuel Torres-Moreno, Miguel Figueroa-Saavedra, Ligia Quintana-Torres, Martha-Lorena Avendaño-Garrido, Graham Ranger

机构 * Laboratoire Informatique d’Avignon(阿维尼翁信息实验室) Avignon Université(阿维尼昂大学) Facultad de Matemáticas(数学系) Instituto de Investigaciones en Educación(教育研究所) Universidad Veracruzana(韦拉克鲁斯大学)

专题命中 预训练与数据 :language model(abstract);分类 cs.CL、cs.AI

Comments 11 pages, 7 tables, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02084 2025-10-06 cs.LG cs.AI 62%

KAIROS: Unified Training for Universal Non-Autoregressive Time Series Forecasting

Kuiye Ding, Fanda Fan, Zheya Wang, Hongxiao Li, Yifan Wang, Lei Wang, Chunjie Luo, Jianfeng Zhan

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Department of Mathematical Sciences, Durham University(杜伦大学数学科学系) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 预训练与数据 :foundation model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00186 2025-10-06 cs.AI cs.LG 62%

Thinkquel: A Model Dedicated to Text-to-dbt Using Synthetic Data and a Span-Aware Objective

Anni Li, Aria Attar, Paul Dong

机构 * TensorStax

专题命中 预训练与数据 :SFT(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02272 2025-10-03 cs.CL cs.AI 62%

Parallel Scaling Law: Unveiling Reasoning Generalization through A Cross-Linguistic Perspective

Wen Yang, Junhong Wu, Chong Li, Chengqing Zong, Jiajun Zhang

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 预训练与数据 :post-training(abstract);分类 cs.CL、cs.AI

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00582 2025-10-02 cs.CL cs.AI cs.SD 62%

SAGE-LD: Towards Scalable and Generalizable End-to-End Language Diarization via Simulated Data Augmentation

Sangmin Lee, Woongjib Choi, Jihyun Kim, Hong-Goo Kang

机构 * Dept. of Electrical \& Electronic Engineering, Yonsei University, Seoul, South Korea

专题命中 预训练与数据 :pretraining(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24988 2025-09-30 cs.CL cs.AI 62%

Generalized Correctness Models: Learning Calibrated and Model-Agnostic Correctness Predictors from Historical Patterns

Hanqi Xiao, Vaidehi Patil, Hyunji Lee, Elias Stengel-Eskin, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 预训练与数据 :LLM(abstract);分类 cs.CL、cs.AI

Comments Code: https://github.com/The-Inscrutable-X/CalibratedModelAgnosticCorrectness

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06246 2025-09-30 cs.CL cs.AI cs.DS 62%

A Partition Cover Approach to Tokenization

Jia Peng Lim, Shawn Tan, Davin Choo, Hady W. Lauw

机构 * Singapore Management University(新加坡管理大学) MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室) Harvard University(哈佛大学) National University of Singapore(新加坡国立大学)

专题命中 预训练与数据 :language model(abstract);分类 cs.CL、cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23261 2025-09-29 cs.CL cs.LG 62%

$100K or 100 Days: Trade-offs when Pre-Training with Academic Resources

Apoorv Khandelwal, Tian Yun, Nihal V. Nayak, Jack Merullo, Stephen H. Bach, Chen Sun, Ellie Pavlick

专题命中 预训练与数据 :pretraining(abstract);分类 cs.CL、cs.LG

Comments Published at COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.00177 2025-09-25 cs.LG cs.AI 62%

Pretrained deep models outperform GBDTs in Learning-To-Rank under label scarcity

Charlie Hou, Kiran Koshy Thekumparampil, Michael Shavlovsky, Giulia Fanti, Yesh Dattatreya, Sujay Sanghavi

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) Carnegie Mellon University(卡内基梅隆大学) Amazon(亚马逊) Department of Computer Science(计算机科学系) University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 预训练与数据 :pretraining(abstract);分类 cs.AI、cs.LG

Comments Published in TMLR, ICML-MFPL 2023 Workshop Oral, SPIGM@ICML2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08653 2025-09-12 cs.LG cs.CL 62%

Generative Data Refinement: Just Ask for Better Data

Minqi Jiang, João G. M. Araújo, Will Ellsworth, Sian Gooding, Edward Grefenstette

机构 * Google DeepMind(谷歌DeepMind)

专题命中 预训练与数据 :prompting(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06888 2025-09-09 cs.CL cs.IR cs.LG 62%

mmBERT: A Modern Multilingual Encoder with Annealed Language Learning

Marc Marone, Orion Weller, William Fleshman, Eugene Yang, Dawn Lawrie, Benjamin Van Durme

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 预训练与数据 :language model(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08228 2025-09-09 cs.LG cs.AI cs.RO 62%

Scaling Laws of Motion Forecasting and Planning -- Technical Report

Mustafa Baniodeh, Kratarth Goel, Scott Ettinger, Carlos Fuertes, Ari Seff, Tim Shen, Cole Gulino, Chenjie Yang, Ghassen Jerfel, Dokook Choe, Rui Wang, Benjamin Charrow, Vinutha Kallem, Sergio Casas, Rami Al-Rfou, Benjamin Sapp, Dragomir Anguelov

机构 * Waymo LLC(Waymo公司)

专题命中 预训练与数据 :language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04465 2025-09-08 cs.CL cs.AI 62%

Emotionally-Aware Agents for Dispute Resolution

Sushrita Rakshit, James Hale, Kushal Chawla, Jeanne M. Brett, Jonathan Gratch

机构 * University of Michigan(密歇根大学) University of Southern California(南加州大学) Northwestern University(西北大学)

专题命中 预训练与数据 :language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22381 2025-09-05 cs.LG cs.AI stat.CO stat.ML 62%

Robust training of implicit generative models for multivariate and heavy-tailed distributions with an invariant statistical loss

José Manuel de Frutos, Manuel A. Vázquez, Pablo Olmos, Joaquín Míguez

机构 * Department of Signal Theory and Communications, Carlos III University of Madrid(信号理论与通信系,卡洛斯三世大学)

专题命中 预训练与数据 :pretraining(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03353 2025-09-04 cs.LG cs.AI 62%

Fair Resource Allocation for Fleet Intelligence

Oguzhan Baser, Kaan Kale, Po-han Li, Sandeep Chinchali

机构 * Wireless Networking Communications Group (WNCG) Department of Electrical Computer Engineering The University of Texas at Austin Email

专题命中 预训练与数据 :language model(abstract);分类 cs.AI、cs.LG

Comments This paper has been accepted for presentation at the 2025 IEEE Global Communications Conference (GLOBECOM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00357 2025-09-03 cs.CV cs.AI cs.LG 62%

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding

Zhen Chen, Xingjian Luo, Kun Yuan, Jinlin Wu, Danny T. M. Chan, Nassir Navab, Hongbin Liu, Zhen Lei, Jiebo Luo

机构 * Hong Kong Institute of Science & Innovation(香港科学与工业创新研究院) CAMP, Technische Universität München(CAMP,慕尼黑技术大学) Department of Surgery, Faculty of Medicine, The Chinese University of Hong Kong(香港中文大学医学院外科部)

专题命中 预训练与数据 :pretraining(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.09513 2025-08-29 cs.NI cs.AI cs.LG 62%

NetGPT: Generative Pretrained Transformer for Network Traffic

Xuying Meng, Chungang Lin, Yequan Wang, Yujun Zhang

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Beijing Academy of Artificial Intelligence(北京人工智能研究院)

专题命中 预训练与数据 :pretraining(abstract);分类 cs.AI、cs.LG

Comments Code is available at https://github.com/ict-net/NetGPT

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24073 2025-08-27 cs.AI cs.CL cs.CV 62%

mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation

Chan-Wei Hu, Yueqi Wang, Shuo Xing, Chia-Ju Chen, Suofei Feng, Ryan Rossi, Zhengzhong Tu

机构 * Texas A&M University(德克萨斯农工大学) University of California, Berkeley(加州大学伯克利分校) University of Texas at Austin(德克萨斯大学奥斯汀分校) Stanford University(斯坦福大学) Adobe Research(Adobe研究院)

专题命中 预训练与数据 :language model(abstract);分类 cs.CL、cs.AI

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17547 2025-08-26 cs.RO cs.AI cs.LG 62%

LodeStar: Long-horizon Dexterity via Synthetic Data Augmentation from Human Demonstrations

Weikang Wan, Jiawei Fu, Xiaodi Yuan, Yifeng Zhu, Hao Su

机构 * University of California San Diego(加州大学圣地亚哥分校) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 预训练与数据 :foundation model(abstract);分类 cs.AI、cs.LG

Comments CoRL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04403 2025-08-25 cs.CL cs.AI 62%

Establishing Task Scaling Laws via Compute-Efficient Model Ladders

Akshita Bhagia, Jiacheng Liu, Alexander Wettig, David Heineman, Oyvind Tafjord, Ananya Harsh Jha, Luca Soldaini, Noah A. Smith, Dirk Groeneveld, Pang Wei Koh, Jesse Dodge, Hannaneh Hajishirzi

机构 * Allen Institute for AI(艾伦人工智能研究所) University of Washington(华盛顿大学) Princeton University(普林斯顿大学)

专题命中 预训练与数据 :language model(abstract);分类 cs.CL、cs.AI

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏