arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 12287 信号源:cs.CL, cs.AI, cs.LG

1. 其他LLM 12287 篇

2506.00432 2025-06-03 cs.LG cs.AI stat.ML 62%

Channel Normalization for Time Series Channel Identification

Seunghan Lee, Taeyoung Park, Kibok Lee

机构 * Yonsei University(延世大学) Department of Statistics and Data Science, Yonsei University(延世大学统计与数据科学系)

专题命中 其他LLM :foundation model(abstract);分类 cs.AI、cs.LG

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13259 2025-06-03 cs.CL cs.AI cs.CY 62%

HumT DumT: Measuring and controlling human-like language in LLMs

Myra Cheng, Sunny Yu, Dan Jurafsky

机构 * Stanford University(斯坦福大学)

专题命中 其他LLM :LLM(abstract);分类 cs.CL、cs.AI

Comments Accepted to ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00659 2025-06-03 cs.LG cs.CL 62%

Why Are Positional Encodings Nonessential for Deep Autoregressive Transformers? Revisiting a Petroglyph

Kazuki Irie

机构 * Harvard University(哈佛大学)

专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.LG

Comments Accepted to ACL 2025 Findings, Short paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22815 2025-06-02 cs.CV cs.AI cs.LG 62%

IMTS is Worth Time $\times$ Channel Patches: Visual Masked Autoencoders for Irregular Multivariate Time Series Prediction

Zhangyi Hu, Jiemin Wu, Hua Xu, Mingqian Liao, Ninghui Feng, Bo Gao, Songning Lai, Yutao Yue

专题命中 其他LLM :foundation model(abstract);分类 cs.AI、cs.LG

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21918 2025-05-29 cs.LG cs.AI 62%

Self-supervised Learning Method Using Transformer for Multi-dimensional Sensor Data Processing

Haruki Kai, Tsuyoshi Okita

专题命中 其他LLM :language model(abstract);分类 cs.AI、cs.LG

Comments 25 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02633 2025-05-29 cs.CR cs.AI cs.LG 62%

Edit Distance Robust Watermarks via Indexing Pseudorandom Codes

Noah Golowich, Ankur Moitra

机构 * MIT(麻省理工学院)

专题命中 其他LLM :language model(abstract);分类 cs.AI、cs.LG

Comments Appeared in NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20482 2025-05-28 cs.CL cs.AI 62%

Conversation Kernels: A Flexible Mechanism to Learn Relevant Context for Online Conversation Understanding

Vibhor Agarwal, Arjoo Gupta, Suparna De, Nishanth Sastry

专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI

Comments Accepted at International AAAI Conference on Web and Social Media (ICWSM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20029 2025-05-27 q-bio.NC cs.AI cs.LG 62%

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)

Subba Reddy Oota, Akshett Jindal, Ishani Mondal, Khushbu Pahwa, Satya Sai Srinath Namburi, Manish Shrivastava, Maneesh Singh, Bapi S. Raju, Manish Gupta

机构 * Technische Universität Berlin(柏林技术大学) IIIT Hyderabad(海得拉巴国家理工学院) Univ of Maryland(马里兰大学) Rice Univ(Rice 大学) Univ of Wisconsin - Madison(威斯康星大学麦迪逊分校) Spector Inc(Spector 公司) Microsoft(微软公司)

专题命中 其他LLM :language model(abstract);分类 cs.AI、cs.LG

Comments 30 pages, 22 figures, The Thirteenth International Conference on Learning Representations, ICLR-2025, Singapore. https://openreview.net/pdf?id=xkgfLXZ4e0

Journal ref ICLR-2025, Singapore

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12632 2025-05-27 cs.CL cs.AI 62%

What External Knowledge is Preferred by LLMs? Characterizing and Exploring Chain of Evidence in Imperfect Context for Multi-Hop QA

Zhiyuan Chang, Mingyang Li, Xiaojun Jia, Junjie Wang, Yuekai Huang, Qing Wang, Yihao Huang, Yang Liu

机构 * State Key Laboratory of Intelligent Game(智能游戏国家重点实验室) Science and Technology on Integrated Information System Laboratory(集成信息系统技术实验室) Institute of Software Chinese Academy of Sciences(中国科学院软件研究所) University of Chinese Academy of Sciences(中国科学院大学) Nanyang Technological University(南洋理工大学)

专题命中 其他LLM :LLM(abstract);分类 cs.CL、cs.AI

Comments 15 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12761 2025-05-26 cs.LG cs.AI 62%

Enhancing Channel-Independent Time Series Forecasting via Cross-Variate Patch Embedding

Donghwa Shin, Edwin Zhang

专题命中 其他LLM :LLM(abstract);分类 cs.AI、cs.LG

Comments Added link to code implementation in PDF abstract

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09286 2025-05-22 cs.RO cs.AI cs.LG 62%

Learning Novel Skills from Language-Generated Demonstrations

Ao-Qun Jin, Tian-Yu Xiang, Xiao-Hu Zhou, Mei-Jiang Gui, Xiao-Liang Xie, Shi-Qi Liu, Shuang-Yi Wang, Yue Cao, Sheng-Bin Duan, Fu-Chao Xie, Zeng-Guang Hou

机构 * Institute of Automation Chinese Academy of Sciences(自动化研究所,中国科学院)

专题命中 其他LLM :language model(abstract);分类 cs.AI、cs.LG

Comments 10 pages, International Conference on Learning Representations (ICLR) 2025 Workshop on Generative Models for Robot Learning (GenBot)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13244 2025-05-20 cs.CL cs.LG 62%

JNLP at SemEval-2025 Task 11: Cross-Lingual Multi-Label Emotion Detection Using Generative Models

Jieying Xue, Phuong Minh Nguyen, Minh Le Nguyen, Xin Liu

机构 * Japan Advanced Institute of Science and Technology(日本先进科学研究所) ROIS-DS Center for Juris-Informatics, NII, Tokyo, Japan(日本信息机构ROIS-DS法律信息中心) National Institute of Advanced Industrial Science and Technology(国家先进工业科学和技术研究所)

专题命中 其他LLM :LLM(abstract);分类 cs.CL、cs.LG

Comments Published in The 19th International Workshop on Semantic Evaluation (SemEval-2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16565 2025-05-20 cs.MA cs.AI cs.CL cs.CY 62%

The Hidden Strength of Disagreement: Unraveling the Consensus-Diversity Tradeoff in Adaptive Multi-Agent Systems

Zengqing Wu, Takayuki Ito

专题命中 其他LLM :LLM(abstract);分类 cs.CL、cs.AI

Comments Source codes are available at https://github.com/wuzengqing001225/ConsensusDiversityTradeoffMAS

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03968 2025-05-20 cs.LG cs.AI cs.GT math.OC 62%

Decoding Game: On Minimax Optimality of Heuristic Text Generation Strategies

Sijin Chen, Omar Hagrass, Jason M. Klusowski

机构 * Department of Electrical and Computer Engineering, Princeton University(普林斯顿大学电气工程与计算机科学系) Department of Operations Research and Financial Engineering, Princeton University(普林斯顿大学运筹学与金融工程系)

专题命中 其他LLM :language model(abstract);分类 cs.AI、cs.LG

Comments 20 pages, accepted to ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08751 2025-05-14 cs.CL cs.CV cs.LG 62%

Aya Vision: Advancing the Frontier of Multilingual Multimodality

Saurabh Dash, Yiyang Nan, John Dang, Arash Ahmadian, Shivalika Singh, Madeline Smith, Bharat Venkitesh, Vlad Shmyhlo, Viraat Aryabumi, Walter Beller-Morales, Jeremy Pekmez, Jason Ozuzu, Pierre Richemond, Acyr Locatelli, Nick Frosst, Phil Blunsom, Aidan Gomez, Ivan Zhang, Marzieh Fadaee, Manoj Govindassamy, Sudip Roy, Matthias Gallé, Beyza Ermis, Ahmet Üstün, Sara Hooker

专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06633 2025-05-13 cs.CL cs.LG 62%

Attention Is Not All You Need: The Importance of Feedforward Networks in Transformer Models

Isaac Gerber

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04223 2025-05-09 cs.LG cs.AI cs.DC 62%

FRAIN to Train: A Fast-and-Reliable Solution for Decentralized Federated Learning

Sanghyeon Park, Soo-Mook Moon

机构 * Seoul National University(首尔国立大学) Theori Inc.(Theori公司)

专题命中 其他LLM :language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02887 2025-05-07 q-bio.BM cs.AI cs.LG 62%

CreoPep: A Universal Deep Learning Framework for Target-Specific Peptide Design and Optimization

Cheng Ge, Han-Shen Tae, Zhenqiang Zhang, Lu Lu, Zhijie Huang, Yilin Wang, Tao Jiang, Wenqing Cai, Shan Chang, David J. Adams, Rilei Yu

机构 * Key Laboratory of Marine Drugs, Chinese Ministry of Education, School of Medicine and Pharmacy, Ocean University of China(海洋大学药理学重点实验室,中国教育部,医学院药学学院,中国海洋大学) Laboratory for Marine Drugs and Bioproducts, Qingdao Marine Science and Technology Center(海洋药物与生物制品实验室,青岛海洋科学与技术中心) Molecular Horizons, Faculty of Science, Medicine and Health, University of Wollongong(分子前沿,科学、医学与健康学院,沃林戈大学) Faculty of Information Science and Engineering, Ocean University of China(信息科学与工程学院,中国海洋大学) Shandong Academy of Pharmaceutical Sciences(山东省药科院) Institute of Bioinformatics and Medical Engineering, School of Electrical and Information Engineering, Jiangsu University of Technology(生物信息学与医学工程学院,江苏科技大学)

专题命中 其他LLM :language model(abstract);分类 cs.AI、cs.LG

Comments 16 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03278 2025-04-16 q-bio.QM cs.AI cs.LG physics.comp-ph 62%

JanusDDG: A Thermodynamics-Compliant Model for Sequence-Based Protein Stability via Two-Fronts Multi-Head Attention

Guido Barducci, Ivan Rossi, Francesco Codicè, Cesare Rollo, Valeria Repetto, Corrado Pancotti, Virginia Iannibelli, Tiziana Sanavia, Piero Fariselli

机构 * University of Turin(都灵大学) Computational Biomedicine Unit, Dept. of Medical Sciences(医学系计算生物医学单元)

专题命中 其他LLM :language model(abstract);分类 cs.AI、cs.LG

Comments 20 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09353 2025-04-16 cs.CV cs.AI cs.CL cs.MM 62%

Causal Graphical Models for Vision-Language Compositional Understanding

Fiorenzo Parascandolo, Nicholas Moratelli, Enver Sangineto, Lorenzo Baraldi, Rita Cucchiara

机构 * University of Modena and Reggio Emilia(摩德纳与雷焦艾米利亚大学) ImageLab(图像实验室)

专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI

Comments Accepted at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08335 2025-04-11 cs.CL cs.LG 62%

Toward a Theory of Tokenization in LLMs

Nived Rajaraman, Jiantao Jiao, Kannan Ramchandran

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.LG

Comments 60 pages, 11 figures. This work was published at NeurIPS 2024 with a different title, "An Analysis of Tokenization: Transformers under Markov data"

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03975 2025-04-08 cs.LG cs.AI 62%

GREATERPROMPT: A Unified, Customizable, and High-Performing Open-Source Toolkit for Prompt Optimization

Wenliang Zheng, Sarkar Snigdha Sarathi Das, Yusen Zhang, Rui Zhang

机构 * Penn State University(宾夕法尼亚州立大学)

专题命中 其他LLM :LLM(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05679 2025-04-08 cs.CV cs.AI cs.LG cs.SD eess.AS 62%

Tell What You Hear From What You See -- Video to Audio Generation Through Text

Xiulong Liu, Kun Su, Eli Shlizerman

机构 * University of Washington(华盛顿大学)

专题命中 其他LLM :LLM(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2024. Project page: https://dragonliu1995.github.io/VATT-home

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01916 2025-04-03 cs.CV cs.AI cs.CL 62%

FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs

Mothilal Asokan, Kebin Wu, Fatima Albreiki

机构 * Technology Innovation Institute (TII)(技术创新研究院)

专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00289 2025-04-03 cs.CV cs.AI cs.LG 62%

Dual Diffusion for Unified Image Generation and Understanding

Zijie Li, Henry Li, Yichun Shi, Amir Barati Farimani, Yuval Kluger, Linjie Yang, Peng Wang

机构 * Carnegie Mellon University(卡内基梅隆大学) Yale University(耶鲁大学) ByteDance(字节跳动)

专题命中 其他LLM :language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23687 2025-04-01 cs.CL cs.LG 62%

MKA: Leveraging Cross-Lingual Consensus for Model Abstention

Sharad Duwal

专题命中 其他LLM :LLM(abstract);分类 cs.CL、cs.LG

Comments To appear in Building Trust Workshop at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22068 2025-03-31 cs.LG cs.AI 62%

A Proposal for Networks Capable of Continual Learning

Zeki Doruk Erden, Boi Faltings

机构 * École Polytechnique Fédérale de Lausanne(洛桑联邦理工学院)

专题命中 其他LLM :prompting(abstract);分类 cs.AI、cs.LG

Comments Published at ICLR 2025 World Models Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02479 2025-03-28 cs.CV cs.AI cs.CR cs.LG 62%

OODFace: Benchmarking Robustness of Face Recognition under Common Corruptions and Appearance Variations

Caixin Kang, Yubo Chen, Shouwei Ruan, Shiji Zhao, Ruochen Zhang, Jiayi Wang, Shan Fu, Xingxing Wei

机构 * Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) China Academy of Information and Communications Technology(中国信息通信研究院)

专题命中 其他LLM :language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20421 2025-03-27 cs.CL cs.LG math.DS 62%

TempTest: Local Normalization Distortion and the Detection of Machine-generated Text

Tom Kempton, Stuart Burrell, Connor Cheverall

机构 * University of Manchester(曼彻斯特大学) Featurespace University of Cambridge(剑桥大学)

专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14714 2025-03-25 cs.AI cs.CL cs.IR 62%

From Knowledge Generation to Knowledge Verification: Examining the BioMedical Generative Capabilities of ChatGPT

Ahmed Abdeen Hamed, Alessandro Crimi, Magdalena M. Misiak, Byung Suk Lee

机构 * MGEN -- College of Engineering, Northeastern University Miami(东北大学迈阿密分校MGEN工程学院) University of Vermont(佛蒙特大学) AGH, University of Krakow(克拉科夫AGH大学) Howard University(霍华德大学)

专题命中 其他LLM :LLM(abstract);分类 cs.CL、cs.AI

Comments 28 pages, 6 figures, In Review with a Cell Press Journal

详情

展开后加载摘要…

URL PDF HTML 收藏