arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 12266 信号源:cs.CL, cs.AI, cs.LG

1. 其他LLM 12266 篇

2506.17144 2025-09-30 cs.CV cs.AI cs.LG 73%

Do We Need Large VLMs for Spotting Soccer Actions?

Ritabrata Chakraborty, Rajatsubhra Chakraborty, Avijit Dasgupta, Sandeep Chaurasia

机构 * Manipal University Jaipur(贾浦尔曼普尔大学) UNC–Charlotte(北卡罗来纳州立大学夏洛特分校) IIIT Hyderabad(海得拉巴印度理工学院)

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

Comments 6 pages, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17550 2025-09-29 cs.LG cs.AI stat.ML 73%

In-Context Algorithm Emulation in Fixed-Weight Transformers

Jerry Yao-Chieh Hu, Hude Liu, Jennifer Yuntong Zhang, Han Liu

专题命中 其他LLM :foundation model(abstract);prompting(abstract);分类 cs.AI、cs.LG

Comments Code is available at https://github.com/MAGICS-LAB/algo_emu

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15561 2025-09-26 cs.LG cs.CL 73%

Small LLMs with Expert Blocks Are Good Enough for Hyperparamter Tuning

Om Naphade, Saksham Bansal, Parikshit Pareek

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02682 2025-09-26 cs.AI cs.LG cs.SY eess.SY math.DS math.OC 73%

The Asymptotic Behavior of Attention in Transformers

Álvaro Rodríguez Abella, João Pedro Silvestre, Paulo Tabuada

机构 * Department of Applied Mathematics, ICAI School of Engineering, Comillas Pontifical University(应用数学系,ICAI工程学院,科利斯大学) Department of Electrical and Computer Engineering, University of California, Los Angeles(电气与计算机工程系,加州大学洛杉矶分校)

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.00016 2025-09-23 q-fin.CP cs.AI cs.CL 73%

Alpha-GPT: Human-AI Interactive Alpha Mining for Quantitative Investment

Saizhuo Wang, Hang Yuan, Leon Zhou, Lionel M. Ni, Heung-Yeung Shum, Jian Guo

机构 * HKUST(香港科技大学) IDEA Research(IDEA研究所) Columbia University(哥伦比亚大学) HKUST-GZ(香港科技大学-广州)

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025 System Demonstration Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16060 2025-09-22 cs.LG cs.CL 73%

SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual Connection

Maithili Joshi, Palash Nandi, Tanmoy Chakraborty

机构 * Indian Institute of Technology Delhi(印度理工学院德里)

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

Comments Accepted in EMNLP'25 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15641 2025-09-22 cs.LG cs.AI stat.ML 73%

Information Geometry of Variational Bayes

Mohammad Emtiyaz Khan

机构 * RIKEN Center for AI Project(日本理化学研究所人工智能项目中心)

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00132 2025-09-18 cs.CL cs.LG 73%

Contextualize-then-Aggregate: Circuits for In-Context Learning in Gemma-2 2B

Aleksandra Bakalova, Yana Veitsman, Xinting Huang, Michael Hahn

机构 * Saarland Informatics Campus(萨尔兰大学信息技术校区)

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11991 2025-09-16 cs.CL cs.AI 73%

Text Adaptation to Plain Language and Easy Read via Automatic Post-Editing Cycles

Jesús Calleja, David Ponce, Thierry Etchegoyhen

机构 * Fundación Vicomtech, Basque Research and Technology Alliance (BRTA)(维克姆科技基金会、巴斯克研究与技术联盟) University of the Basque Country - UPV/EHU(巴斯克国家大学 - UPV/EHU)

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10931 2025-09-16 cs.AI cs.CL 73%

Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding

Seongho Joo, Hyukhun Koh, Kyomin Jung

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09713 2025-09-15 cs.CL cs.AI 73%

HANRAG: Heuristic Accurate Noise-resistant Retrieval-Augmented Generation for Multi-hop Question Answering

Duolin Sun, Dan Yang, Yue Shen, Yihan Jiao, Zhehao Tan, Jie Feng, Lianzhen Zhong, Jian Wang, Peng Wei, Jinjie Gu

机构 * Ant Group(蚂蚁集团)

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07226 2025-09-10 astro-ph.EP astro-ph.IM cs.AI cs.LG 73%

A transformer-based generative model for planetary systems

Yann Alibert, Jeanne Davoult, Sara Marques

机构 * Space Research and Planetary Sciences, Physics Institute, University of Bern(空间研究与行星科学系,伯尔尼大学) Center for Space and Habitability, University of Bern(空间与宜居性研究中心,伯尔尼大学) Institut für Planetenforschung, German Aerospace Center (DLR)(行星研究研究所,德国航空航天中心(DLR))

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

Comments Accepted in A&A

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12761 2025-09-09 cs.CL cs.AI 73%

Revealing the impact of synthetic native samples and multi-tasking strategies in Hindi-English code-mixed humour and sarcasm detection

Debajyoti Mazumder, Aakash Kumar, Jasabanta Patro

机构 * Department of Data Science and Engineering(数据科学与工程系) Indian Institute of Science Education and Research(印度科学教育与研究学院)

专题命中 其他LLM :language model(abstract);prompting(abstract);分类 cs.CL、cs.AI

Comments 33 pages; EMNLP 2025 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04479 2025-09-08 cs.CL cs.AI 73%

No Clustering, No Routing: How Transformers Actually Process Rare Tokens

Jing Liu

机构 * ENS, Université PSL, EHESS, CNRS Paris, France(巴黎大学PSL研究所、EHESS、CNRS)

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01684 2025-09-04 cs.LG cs.AI 73%

Reinforcement Learning for Machine Learning Engineering Agents

Sherry Yang, Joy He-Yueya, Percy Liang

机构 * Stanford University(斯坦福大学)

专题命中 其他LLM :language model(abstract);prompting(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00949 2025-09-03 cs.CL cs.AI 73%

Structure and Destructure: Dual Forces in the Making of Knowledge Engines

Yihong Chen

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments PhD thesis. https://discovery.ucl.ac.uk/id/eprint/10211291/

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17625 2025-09-01 cs.LG cs.AI 73%

Alice's Adventures in a Differentiable Wonderland -- Volume I, A Tour of the Land

Simone Scardapane

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

Comments Companion website for additional chapters: https://www.sscardapane.it/alice-book

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19427 2025-08-28 cs.CL cs.AI cs.CY cs.HC 73%

A perishable ability? The future of writing in the face of generative artificial intelligence

Evandro L. T. P. Cunha

机构 * Universidade Federal de Minas Gerais (UFMG)(巴西联邦大学矿务学院)

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18724 2025-08-27 cs.AI cs.CL 73%

Bias Mitigation Agent: Optimizing Source Selection for Fair and Balanced Knowledge Retrieval

Karanbir Singh, Deepak Muppiri, William Ngu

机构 * Salesforce

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments Accepted at KDD'2025 Agent4IR workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13810 2025-08-27 cs.CV cs.AI cs.LG cs.RO 73%

CAD-Assistant: Tool-Augmented VLLMs as Generic CAD Task Solvers

Dimitrios Mallis, Ahmet Serdar Karadeniz, Sebastian Cavada, Danila Rukhovich, Niki Foteinopoulou, Kseniya Cherenkova, Anis Kacem, Djamila Aouada

机构 * SnT, University of Luxembourg(斯诺特研究所,卢森堡大学) Artec3D, Luxembourg(Artec3D,卢森堡)

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06866 2025-08-27 cs.CL cs.AI 73%

ChatGPT Doesn't Trust Chargers Fans: Guardrail Sensitivity in Context

Victoria R. Li, Yida Chen, Naomi Saphra

机构 * Harvard University(哈佛大学) Kempner Institute for the Study of Natural and Artificial Intelligence(凯普纳人工智能研究 institute)

专题命中 其他LLM :LLM(abstract);language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11829 2025-08-19 cs.CL cs.AI cs.MA 73%

Every 28 Days the AI Dreams of Soft Skin and Burning Stars: Scaffolding AI Agents with Hormones and Emotions

Leigh Levinson, Christopher J. Agostino

机构 * Indiana University(印第安纳大学) NPC Worldwide(NPC全球)

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments 9 pages, 1 figure, submitted to NeurIPS Creative AI track

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08052 2025-08-15 cs.LG cs.AI 73%

On Understanding of the Dynamics of Model Capacity in Continual Learning

Supriyo Chakraborty, Krishnan Raghavan

机构 * AI Foundations Capital One(AI基础资本一公司) Mathematics and Computer Science(数学与计算机科学) Argonne National Laboratory(阿贡国家实验室)

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14937 2025-08-05 cs.LG cs.AI cs.CR 73%

Geminio: Language-Guided Gradient Inversion Attacks in Federated Learning

Junjie Shan, Ziqi Zhao, Jialin Lu, Rui Zhang, Siu Ming Yiu, Ka-Ho Chow

机构 * School of Computing and Data Science(计算与数据科学学院) The University of Hong Kong(香港大学)

专题命中 其他LLM :language model(abstract);foundation model(abstract);分类 cs.AI、cs.LG

Comments ICCV2025 camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00760 2025-08-04 cs.CL cs.AI 73%

MMBERT: Scaled Mixture-of-Experts Multimodal BERT for Robust Chinese Hate Speech Detection under Cloaking Perturbations

Qiyao Xue, Yuchen Dou, Ryan Shi, Xiang Lorraine Li, Wei Gao

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16604 2025-08-04 cs.CL cs.AI cs.SI 73%

Debunking with Dialogue? Exploring AI-Generated Counterspeech to Challenge Conspiracy Theories

Mareike Lisker, Christina Gottschalk, Helena Mihaljević

机构 * University of Applied Sciences (HTW) Berlin, Germany(柏林应用科学大学(HTW))

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments 16 pages, Association for Computational Linguistics, Proceedings of the 9th Workshop on Online Abuse and Harms (WOAH 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.13481 2025-08-01 cs.CY cs.AI cs.CL cs.HC 73%

How AI Ideas Affect the Creativity, Diversity, and Evolution of Human Ideas: Evidence From a Large, Dynamic Experiment

Joshua Ashkinaze, Julia Mendelsohn, Li Qiwei, Ceren Budak, Eric Gilbert

机构 * University of Michigan(密歇根大学) University of Maryland(马里兰大学)

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments Accepted at ACM Collective Intelligence 2025. Originally posted 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16991 2025-07-29 cs.LG cs.AI 73%

PyG 2.0: Scalable Learning on Real World Graphs

Matthias Fey, Jinu Sunil, Akihiro Nitta, Rishi Puri, Manan Shah, Blaž Stojanovič, Ramona Bendias, Alexandria Barghi, Vid Kocijan, Zecheng Zhang, Xinwei He, Jan Eric Lenssen, Jure Leskovec

机构 * Nvidia Max Planck Institute for Informatics(马克斯·普朗克研究所) Stanford University(斯坦福大学)

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18202 2025-07-25 cs.CL cs.AI 73%

Safeguarding RAG Pipelines with GMTP: A Gradient-based Masked Token Probability Method for Poisoned Document Detection

San Kim, Jonghwi Kim, Yejin Jeon, Gary Geunbae Lee

机构 * Graduate School of Artificial Intelligence, POSTECH, Republic of Korea(人工智能研究生院,POSTECH,韩国) Department of Computer Science and Engineering, POSTECH, Republic of Korea(计算机科学与工程系,POSTECH,韩国)

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments 18 pages, accepted to ACL Findings 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06565 2025-07-24 cs.CL cs.LG 73%

A Mathematical Theory of Discursive Networks

Juan B. Gutiérrez

机构 * Department of Mathematics, University of Texas at San Antonio(德克萨斯大学圣安东尼奥分校数学系)

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

Comments 42 pages, 4 figures, 4 tables, 3 algorithm, 61 references

详情

展开后加载摘要…

URL PDF HTML 收藏