arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 139420 信号源:cs.CL, cs.AI, cs.LG

1. 预训练与数据 12429 篇

2402.16024 2024-05-21 cs.CL cs.LG 76%

HiGPT: Heterogeneous Graph Language Model

Jiabin Tang, Yuhao Yang, Wei Wei, Lei Shi, Long Xia, Dawei Yin, Chao Huang

专题命中 预训练与数据 :language model(title);分类 cs.CL、cs.LG

Comments Accepted by KDD'2024, full paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10714 2024-05-20 cs.CL cs.AI 76%

Persian Pronoun Resolution: Leveraging Neural Networks and Language Models

Hassan Haji Mohammadi, Alireza Talebpour, Ahmad Mahmoudi Aznaveh, Samaneh Yazdani

专题命中 预训练与数据 :language model(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11984 2024-03-19 cs.CL cs.AI cs.HC 76%

Using Generative Text Models to Create Qualitative Codebooks for Student Evaluations of Teaching

Andrew Katz, Mitchell Gerhardt, Michelle Soledad

专题命中 预训练与数据 :large language model(abstract,comments);language model(abstract,comments);分类 cs.CL、cs.AI

Comments Natural language processing, large language models, generative AI, student evaluations of teaching, codebook generation, qualitative data analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.16454 2024-01-31 cs.HC cs.AI cs.CL cs.IR 76%

KAUCUS: Knowledge Augmented User Simulators for Training Language Model Assistants

Kaustubh D. Dhole

专题命中 预训练与数据 :language model(title);分类 cs.CL、cs.AI

Comments Simulation of Conversational Intelligence in Chat, EACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.05570 2024-01-12 cs.CV cs.AI cs.LG 76%

Siamese Networks with Soft Labels for Unsupervised Lesion Detection and Patch Pretraining on Screening Mammograms

Kevin Van Vorst, Li Shen

专题命中 预训练与数据 :pretraining(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.08789 2023-11-10 cs.RO cs.AI cs.LG 76%

PLEX: Making the Most of the Available Data for Robotic Manipulation Pretraining

Garrett Thomas, Ching-An Cheng, Ricky Loynd, Felipe Vieira Frujeri, Vibhav Vineet, Mihai Jalobeanu, Andrey Kolobov

专题命中 预训练与数据 :pretraining(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08304 2023-10-13 cs.CV cs.AI cs.LG 76%

CHIP: Contrastive Hierarchical Image Pretraining

Arpit Mittal, Harshil Jhaveri, Swapnil Mallick, Abhishek Ajmera

专题命中 预训练与数据 :pretraining(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04750 2023-10-11 cs.AI cs.CV cs.LG 76%

DiffNAS: Bootstrapping Diffusion Models by Prompting for Better Architectures

Wenhao Li, Xiu Su, Shan You, Fei Wang, Chen Qian, Chang Xu

专题命中 预训练与数据 :prompting(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12898 2023-08-28 cs.MM cs.AI cs.CL cs.CV 76%

Can Linguistic Knowledge Improve Multimodal Alignment in Vision-Language Pretraining?

Fei Wang, Liang Ding, Jun Rao, Ye Liu, Li Shen, Changxing Ding

专题命中 预训练与数据 :pretraining(title);分类 cs.CL、cs.AI

Comments [TL;DR] we design and release the SNARE, the first large-scale multimodal alignment probing benchmark for current vision-language pretrained models

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.13947 2023-06-27 cs.CL cs.LG 76%

Comparison of Pre-trained Language Models for Turkish Address Parsing

Muhammed Cihat Ünal, Betül Aygün, Aydın Gerek

专题命中 预训练与数据 :language model(title);分类 cs.CL、cs.LG

Comments published in 16th UYMS (2023) https://ekitap.atauni.edu.tr/index.php/product/16-ulusal-yazilim-muhendisligi-sempozyumu-bildiri-kitabi/

Journal ref Ulusal Yazılım Mühendisliği Sempozyumu, 16, 155-165 (Erzurum 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.12561 2023-06-07 cs.CV cs.CL cs.LG 76%

Retrieval-Augmented Multimodal Language Modeling

Michihiro Yasunaga, Armen Aghajanyan, Weijia Shi, Rich James, Jure Leskovec, Percy Liang, Mike Lewis, Luke Zettlemoyer, Wen-tau Yih

专题命中 预训练与数据 :language model(title);分类 cs.CL、cs.LG

Comments Published at ICML 2023. Blog post available at https://cs.stanford.edu/~myasu/blog/racm3/

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.02153 2023-06-06 cs.CL cs.LG cs.SD eess.AS 76%

Acoustic Word Embeddings for Untranscribed Target Languages with Continued Pretraining and Learned Pooling

Ramon Sanabria, Ondrej Klejch, Hao Tang, Sharon Goldwater

专题命中 预训练与数据 :pretraining(title);分类 cs.CL、cs.LG

Comments Accepted to Interspeech 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.08356 2023-06-01 cs.CV cs.AI cs.LG stat.ML 76%

OmniMAE: Single Model Masked Pretraining on Images and Videos

Rohit Girdhar, Alaaeldin El-Nouby, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, Ishan Misra

专题命中 预训练与数据 :pretraining(title);分类 cs.AI、cs.LG

Comments CVPR 2023. Code/models: https://github.com/facebookresearch/omnivore

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.14396 2023-03-28 cs.CV cs.AI cs.LG 76%

IFSeg: Image-free Semantic Segmentation via Vision-Language Model

Sukmin Yun, Seong Hyeon Park, Paul Hongsuck Seo, Jinwoo Shin

专题命中 预训练与数据 :language model(title);分类 cs.AI、cs.LG

Comments Accepted to CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.01893 2023-01-06 cs.CV cs.AI cs.CL 76%

GIVL: Improving Geographical Inclusivity of Vision-Language Models with Pre-Training Methods

Da Yin, Feng Gao, Govind Thattai, Michael Johnston, Kai-Wei Chang

专题命中 预训练与数据 :language model(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.07036 2022-11-29 cs.CL cs.AI cs.CV cs.SD eess.AS eess.IV 76%

u-HuBERT: Unified Mixed-Modal Speech Pretraining And Zero-Shot Transfer to Unlabeled Modality

Wei-Ning Hsu, Bowen Shi

专题命中 预训练与数据 :pretraining(title);分类 cs.CL、cs.AI

Comments NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.07889 2022-11-16 cs.LG cs.AI eess.SP 76%

Pretraining ECG Data with Adversarial Masking Improves Model Generalizability for Data-Scarce Tasks

Jessica Y. Bo, Hen-Wei Huang, Alvin Chan, Giovanni Traverso

专题命中 预训练与数据 :pretraining(title);分类 cs.AI、cs.LG

Comments Extended Abstract presented at Machine Learning for Health (ML4H) symposium 2022, November 28th, 2022, New Orleans, United States & Virtual, http://www.ml4h.cc, 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.10913 2022-10-25 cs.LG cs.AI 76%

Palm up: Playing in the Latent Manifold for Unsupervised Pretraining

Hao Liu, Tom Zahavy, Volodymyr Mnih, Satinder Singh

专题命中 预训练与数据 :pretraining(title);分类 cs.AI、cs.LG

Comments Thirty-sixth Conference on Neural Information Processing Systems (NeurIPS 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.10049 2022-07-21 cs.CV cs.AI cs.LG 76%

Pretraining a Neural Network before Knowing Its Architecture

Boris Knyazev

专题命中 预训练与数据 :pretraining(title);分类 cs.AI、cs.LG

Comments Accepted at ICML 2022 Workshop on Pre-training: Perspectives, Pitfalls, and Paths Forward, source code is available at https://github.com/facebookresearch/ppuda

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.02272 2022-07-07 cs.CL cs.AI 76%

Pretraining on Interactions for Learning Grounded Affordance Representations

Jack Merullo, Dylan Ebert, Carsten Eickhoff, Ellie Pavlick

专题命中 预训练与数据 :pretraining(title);分类 cs.CL、cs.AI

Comments *SEM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.15868 2022-06-01 cs.CV cs.CL cs.LG 76%

CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, Jie Tang

专题命中 预训练与数据 :pretraining(title);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.07911 2022-04-21 cs.CL cs.LG 76%

Signal in Noise: Exploring Meaning Encoded in Random Character Sequences with Character-Aware Language Models

Mark Chu, Bhargav Srinivasa Desikan, Ethan O. Nadler, D. Ruggiero Lo Sardo, Elise Darragh-Ford, Douglas Guilbeault

专题命中 预训练与数据 :language model(title);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.07464 2022-04-18 cs.CL cs.AI 76%

Improving Pre-trained Language Models with Syntactic Dependency Prediction Task for Chinese Semantic Error Recognition

Bo Sun, Baoxin Wang, Wanxiang Che, Dayong Wu, Zhigang Chen, Ting Liu

专题命中 预训练与数据 :language model(title);分类 cs.CL、cs.AI

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.03542 2022-04-08 cs.CL cs.AI 76%

Leveraging pre-trained language models for conversational information seeking from text

Patrizio Bellan, Mauro Dragoni, Chiara Ghidini

专题命中 预训练与数据 :language model(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.05848 2021-10-19 cs.CL cs.AI 76%

Family of Origin and Family of Choice: Massively Parallel Lexiconized Iterative Pretraining for Severely Low Resource Machine Translation

Zhong Zhou, Alex Waibel

专题命中 预训练与数据 :pretraining(title);分类 cs.CL、cs.AI

Journal ref In Proceedings of the 3rd Workshop on Research in Computational Typology and Multilingual NLP of the 20th Conference of the North American Chapter of the Association for Computational Linguistics on Human Language Technologies in 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.06758 2021-04-13 cs.CL cs.AI 76%

ENTRUST: Argument Reframing with Language Models and Entailment

Tuhin Chakrabarty, Christopher Hidey, Smaranda Muresan

专题命中 预训练与数据 :language model(title);分类 cs.CL、cs.AI

Comments NAACL 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.11363 2021-01-28 cs.CL cs.LG 76%

KoreALBERT: Pretraining a Lite BERT Model for Korean Language Understanding

Hyunjae Lee, Jaewoong Yoon, Bonggyu Hwang, Seongho Joe, Seungjai Min, Youngjune Gwon

专题命中 预训练与数据 :pretraining(title);分类 cs.CL、cs.LG

Comments 7 pages, 1 figure, to be published in 25th International Conference on Pattern Recognition, ICPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.06838 2019-07-17 cs.LG cs.AI cs.CV eess.IV 76%

Improved Reinforcement Learning through Imitation Learning Pretraining Towards Image-based Autonomous Driving

Tianqi Wang, Dong Eui Chang

专题命中 预训练与数据 :pretraining(title);分类 cs.AI、cs.LG

Comments 5 pages, 2019 19th International Conference on Control, Automation and Systems (ICCAS 2019)

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.06503 2018-05-18 cs.CL cs.AI 76%

Weight Initialization in Neural Language Models

Ameet Deshpande, Vedant Somani

专题命中 预训练与数据 :language model(title);分类 cs.CL、cs.AI

Comments 17 pages, 20 figures and/or tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10465 2026-08-11 cs.LG cs.AI cs.CL 75%

Superposition Yields Robust Neural Scaling

叠加产生稳健的神经扩展

Yizhou Liu, Ziming Liu, Jeff Gore

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 预训练与数据 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究发现表示叠加是神经扩展定律的核心驱动因素,揭示了损失与模型规模之间的反比关系。

Comments Best Paper Runner-up at NeurIPS 2025

Journal ref Advances in Neural Information Processing Systems 38 (2025) 159269--159305

详情

展开后加载摘要…

URL PDF HTML 收藏