arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 12502 信号源:cs.CL, cs.AI, cs.LG

1. 预训练与数据 12502 篇

2411.00392 2024-11-04 cs.LG cs.AI 62%

Preventing Dimensional Collapse in Self-Supervised Learning via Orthogonality Regularization

Junlin He, Jinxiao Du, Wei Ma

机构 * The Hong Kong Polytechnic University(香港理工大学)

专题命中 预训练与数据 :pretraining(abstract);分类 cs.AI、cs.LG

Comments accepted by NeurIPS 2024 as a poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04640 2024-11-01 cs.RO cs.AI cs.LG 62%

Unpacking Failure Modes of Generative Policies: Runtime Monitoring of Consistency and Progress

Christopher Agia, Rohan Sinha, Jingyun Yang, Zi-ang Cao, Rika Antonova, Marco Pavone, Jeannette Bohg

机构 * Stanford University(斯坦福大学) NVIDIA Research(英伟达研究院)

专题命中 预训练与数据 :language model(abstract);分类 cs.AI、cs.LG

Comments Project page: https://sites.google.com/stanford.edu/sentinel. 35 pages, 9 figures. Accepted to the Conference on Robot Learning (CoRL) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.04459 2024-11-01 cs.LG cs.AI cs.RO 62%

Subwords as Skills: Tokenization for Sparse-Reward Reinforcement Learning

David Yunis, Justin Jung, Falcon Dai, Matthew Walter

机构 * TTI-Chicago(芝加哥TTI学院) University of Chicago(芝加哥大学) Toyota Technological Institute at Chicago(芝加哥丰田理工学院) Symbolica AI(Symbolica AI公司)

专题命中 预训练与数据 :pretraining(abstract);分类 cs.AI、cs.LG

Comments Accepted to NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23356 2024-11-01 cs.LG cs.AI stat.ML 62%

Sequential Order-Robust Mamba for Time Series Forecasting

Seunghan Lee, Juri Hong, Kibok Lee, Taeyoung Park

机构 * Yonsei University(延世大学)

专题命中 预训练与数据 :pretraining(abstract);分类 cs.AI、cs.LG

Comments NeurIPS Workshop on Time Series in the Age of Large Models, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22803 2024-10-31 cs.SD cs.AI cs.LG cs.MM eess.AS 62%

DOA-Aware Audio-Visual Self-Supervised Learning for Sound Event Localization and Detection

Yoto Fujita, Yoshiaki Bando, Keisuke Imoto, Masaki Onishi, Kazuyoshi Yoshii

机构 * Graduate School of Informatics, Kyoto University(京都大学情报学研究科) National Institute of Advanced Industrial Science and Technology(独立行政法人产业技术综合研究所) Faculty of Science and Engineering, Doshisha University(同志社大学理工学部)

专题命中 预训练与数据 :pretraining(abstract);分类 cs.AI、cs.LG

Comments Accepted to APSIPA2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08461 2024-10-31 cs.LG cs.AI 62%

OTTER: Effortless Label Distribution Adaptation of Zero-shot Models

Changho Shin, Jitian Zhao, Sonia Cromp, Harit Vishwakarma, Frederic Sala

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 预训练与数据 :pretraining(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13882 2024-10-30 cs.LG cs.AI 62%

Tabular Data Generation using Binary Diffusion

Vitaliy Kinakh, Slava Voloshynovskiy

专题命中 预训练与数据 :pretraining(abstract);分类 cs.AI、cs.LG

Comments Accepted to 3rd Table Representation Learning Workshop @ NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00048 2024-10-30 cs.CL cond-mat.dis-nn cs.LG 62%

Towards a theory of how the structure of language is acquired by deep neural networks

Francesco Cagnetta, Matthieu Wyart

专题命中 预训练与数据 :language model(abstract);分类 cs.CL、cs.LG

Comments NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10583 2024-10-30 cs.CL cs.LG 62%

Who Are All The Stochastic Parrots Imitating? They Should Tell Us!

Sagi Shaier, Lawrence E. Hunter, Katharina von der Wense

专题命中 预训练与数据 :language model(abstract);分类 cs.CL、cs.LG

Comments Accepted to IJCNLP-AACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16845 2024-10-29 cs.LG cs.CL stat.ML 62%

On Mesa-Optimization in Autoregressively Trained Transformers: Emergence and Capability

Chenyu Zheng, Wei Huang, Rongzhen Wang, Guoqiang Wu, Jun Zhu, Chongxuan Li

专题命中 预训练与数据 :pretraining(abstract);分类 cs.CL、cs.LG

Comments Accepted by NeurIPS2024, 45 pages. The final version of the previous preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18441 2024-10-25 cs.LG cs.AI 62%

The Nature of Mathematical Modeling and Probabilistic Optimization Engineering in Generative AI

Fulu Li

专题命中 预训练与数据 :language model(abstract);分类 cs.AI、cs.LG

Comments 19 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12819 2024-10-24 cs.DC cs.CL cs.LG cs.NI 62%

I've Got 99 Problems But FLOPS Ain't One

Alexandru M. Gherghescu, Vlad-Andrei Bădoiu, Alexandru Agache, Mihai-Valentin Dumitru, Iuliu Vasilescu, Radu Mantu, Costin Raiciu

专题命中 预训练与数据 :language model(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17094 2024-10-23 cs.CL cs.AI 62%

Team Ryu's Submission to SIGMORPHON 2024 Shared Task on Subword Tokenization

Zilong Li

专题命中 预训练与数据 :language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.08672 2024-10-21 eess.SP cs.AI cs.CV cs.LG 62%

RRWaveNet: A Compact End-to-End Multi-Scale Residual CNN for Robust PPG Respiratory Rate Estimation

Pongpanut Osathitporn, Guntitat Sawadwuthikul, Punnawish Thuwajit, Kawisara Ueafuea, Thee Mateepithaktham, Narin Kunaseth, Tanut Choksatchawathi, Proadpran Punyabukkana, Emmanuel Mignot, Theerawit Wilaiprasitporn

专题命中 预训练与数据 :pretraining(abstract);分类 cs.AI、cs.LG

Comments 11 pages, 8 figures

Journal ref 2023 IEEE Internet of Things Journal, vol. 10, no. 18, pp. 15943-15952

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13045 2024-10-18 cs.LG cs.AI 62%

FedGTST: Boosting Global Transferability of Federated Models via Statistics Tuning

Evelyn Ma, Chao Pan, Rasoul Etesami, Han Zhao, Olgica Milenkovic

专题命中 预训练与数据 :pretraining(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09838 2024-10-17 cs.LG cs.AI cs.CR 62%

Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor Defense

Rui Min, Zeyu Qin, Nevin L. Zhang, Li Shen, Minhao Cheng

专题命中 预训练与数据 :pretraining(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2024 Spotlight paper. The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.02623 2024-10-15 cs.CY cs.AI cs.CL cs.CV 62%

Uplifting Lower-Income Data: Strategies for Socioeconomic Perspective Shifts in Large Multi-modal Models

Joan Nwatu, Oana Ignat, Rada Mihalcea

专题命中 预训练与数据 :prompting(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.20086 2024-10-14 cs.CL cs.LG 62%

Token Erasure as a Footprint of Implicit Vocabulary Items in LLMs

Sheridan Feucht, David Atkinson, Byron Wallace, David Bau

专题命中 预训练与数据 :LLM(abstract);分类 cs.CL、cs.LG

Comments 13 pages, 14 figures. Code and data at https://footprints.baulab.info/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08320 2024-10-14 cs.CL cs.LG 62%

Do You Know What You Are Talking About? Characterizing Query-Knowledge Relevance For Reliable Retrieval Augmented Generation

Zhuohang Li, Jiaxin Zhang, Chao Yan, Kamalika Das, Sricharan Kumar, Murat Kantarcioglu, Bradley A. Malin

专题命中 预训练与数据 :language model(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07472 2024-10-11 cs.LG cs.AI 62%

Exploring the design space of deep-learning-based weather forecasting systems

Shoaib Ahmed Siddiqui, Jean Kossaifi, Boris Bonev, Christopher Choy, Jan Kautz, David Krueger, Kamyar Azizzadenesheli

专题命中 预训练与数据 :pretraining(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07379 2024-10-11 eess.AS cs.AI cs.CL 62%

Learn from Real: Reality Defender's Submission to ASVspoof5 Challenge

Yi Zhu, Chirag Goel, Surya Koppisetti, Trang Tran, Ankur Kumar, Gaurav Bharaj

专题命中 预训练与数据 :pretraining(abstract);分类 cs.CL、cs.AI

Comments Accepted into ASVspoof5 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18164 2024-10-10 cs.CL cs.LG 62%

Nebula: A discourse aware Minecraft Builder

Akshay Chaturvedi, Kate Thompson, Nicholas Asher

专题命中 预训练与数据 :LLM(abstract);分类 cs.CL、cs.LG

Comments EMNLP 2024 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01851 2024-10-08 cs.LG cs.AI cs.PF 62%

The Power of Training: How Different Neural Network Setups Influence the Energy Demand

Daniel Geißler, Bo Zhou, Mengxi Liu, Sungho Suh, Paul Lukowicz

专题命中 预训练与数据 :pretraining(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00999 2024-10-07 cs.LG cs.CL cs.CR 62%

Seeing the Forest through the Trees: Data Leakage from Partial Transformer Gradients

Weijun Li, Qiongkai Xu, Mark Dras

专题命中 预训练与数据 :language model(abstract);分类 cs.CL、cs.LG

Comments accepted to EMNLP2024 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06648 2024-10-07 cs.CL cs.AI cs.IR 62%

Dense X Retrieval: What Retrieval Granularity Should We Use?

Tong Chen, Hongwei Wang, Sihao Chen, Wenhao Yu, Kaixin Ma, Xinran Zhao, Hongming Zhang, Dong Yu

专题命中 预训练与数据 :language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02613 2024-10-04 cs.CV cs.AI cs.CL 62%

NL-Eye: Abductive NLI for Images

Mor Ventura, Michael Toker, Nitay Calderon, Zorik Gekhman, Yonatan Bitton, Roi Reichart

专题命中 预训练与数据 :language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02343 2024-10-04 cs.CL cs.LG 62%

Listening to the Wise Few: Select-and-Copy Attention Heads for Multiple-Choice QA

Eduard Tulchinskii, Laida Kushnareva, Kristian Kuznetsov, Anastasia Voznyuk, Andrei Andriiainen, Irina Piontkovskaya, Evgeny Burnaev, Serguei Barannikov

专题命中 预训练与数据 :LLM(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.07840 2024-10-04 cs.CL cs.LG 62%

On Training Data Influence of GPT Models

Yekun Chai, Qingyi Liu, Shuohuan Wang, Yu Sun, Qiwei Peng, Hua Wu

专题命中 预训练与数据 :language model(abstract);分类 cs.CL、cs.LG

Comments EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14857 2024-10-03 cs.CV cs.AI cs.LG 62%

Conditional Diffusion on Web-Scale Image Pairs leads to Diverse Image Variations

Manoj Kumar, Neil Houlsby, Emiel Hoogeboom

专题命中 预训练与数据 :pretraining(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19541 2024-10-03 cs.CL cs.AI 62%

Unlabeled Debiasing in Downstream Tasks via Class-wise Low Variance Regularization

Shahed Masoudian, Markus Frohmann, Navid Rekabsaz, Markus Schedl

专题命中 预训练与数据 :language model(abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏