arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 12228 信号源:cs.CL, cs.AI, cs.LG

1. 其他LLM 12228 篇

2210.00660 2023-02-08 cs.LG cs.AI cs.CL 82%

A Non-monotonic Self-terminating Language Model

Eugene Choi, Kyunghyun Cho, Cheolhyoung Lee

专题命中 其他LLM :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Published as a conference paper at ICLR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.15029 2022-12-02 cs.CL cs.AI cs.LG 82%

DiffusionBERT: Improving Generative Masked Language Models with Diffusion Models

Zhengfu He, Tianxiang Sun, Kuanning Wang, Xuanjing Huang, Xipeng Qiu

专题命中 其他LLM :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Work in progress. Code publicly available at https://github.com/Hzfinfdu/Diffusion-BERT

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12763 2022-10-25 cs.CL cs.AI cs.LG 82%

Discriminative Language Model as Semantic Consistency Scorer for Prompt-based Few-Shot Text Classification

Zhipeng Xie, Yahe Li

专题命中 其他LLM :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 21 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.02329 2022-10-11 cs.CL cs.AI cs.LG 82%

Can language models learn from explanations in context?

Andrew K. Lampinen, Ishita Dasgupta, Stephanie C. Y. Chan, Kory Matthewson, Michael Henry Tessler, Antonia Creswell, James L. McClelland, Jane X. Wang, Felix Hill

专题命中 其他LLM :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Findings of EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.14554 2022-10-05 cs.CL cs.AI cs.IT cs.LG math.IT 82%

Do language models make human-like predictions about the coreferents of Italian anaphoric zero pronouns?

James A. Michaelov, Benjamin K. Bergen

专题命中 其他LLM :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at COLING 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.11902 2022-09-27 cs.AI cs.CL cs.GT cs.LG 82%

Learning Chess With Language Models and Transformers

Michael DeLeo, Erhan Guven

专题命中 其他LLM :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Conference Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.00795 2022-02-03 cs.CL cs.AI cs.LG 82%

Disaster Tweets Classification using BERT-Based Language Model

Anh Duc Le

专题命中 其他LLM :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments arXiv admin note: text overlap with arXiv:2102.12162 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.13654 2021-11-29 cs.CL cs.AI cs.LG 82%

Do Language Models Have Beliefs? Methods for Detecting, Updating, and Visualizing Model Beliefs

Peter Hase, Mona Diab, Asli Celikyilmaz, Xian Li, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, Srinivasan Iyer

专题命中 其他LLM :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.11715 2021-09-14 cs.CL cs.AI cs.LG cs.NE cs.SD eess.AS 82%

Multi-task Language Modeling for Improving Speech Recognition of Rare Words

Chao-Han Huck Yang, Linda Liu, Ankur Gandhe, Yile Gu, Anirudh Raju, Denis Filimonov, Ivan Bulyko

专题命中 其他LLM :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to IEEE Automatic Speech Recognition and Understanding (ASRU) 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.00857 2021-07-13 cs.CL cs.AI cs.LG 82%

StructFormer: Joint Unsupervised Induction of Dependency and Constituency Structure from Masked Language Modeling

Yikang Shen, Yi Tay, Che Zheng, Dara Bahri, Donald Metzler, Aaron Courville

专题命中 其他LLM :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Published as a conference paper at ACL 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.00737 2020-02-04 cs.CL cs.AI cs.LG 82%

Are Pre-trained Language Models Aware of Phrases? Simple but Strong Baselines for Grammar Induction

Taeuk Kim, Jihun Choi, Daniel Edmiston, Sang-goo Lee

专题命中 其他LLM :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments ICLR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.11503 2019-10-28 cs.SE 82%

Improve Language Modelling for Code Completion through Statement Level Language Model based on Statement Embedding Generated by BiLSTM

Yixiao Yang

专题命中 其他LLM :language model(title,abstract);SLM(abstract)

Comments The experimental data is not complete and has some error!

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.10705 2019-09-25 cs.CL cs.AI cs.LG 82%

Do Massively Pretrained Language Models Make Better Storytellers?

Abigail See, Aneesh Pappu, Rohun Saxena, Akhila Yerukola, Christopher D. Manning

专题命中 其他LLM :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to CoNLL 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.04458 2019-06-25 cs.CL cs.AI cs.LG 82%

Knowledge-Augmented Language Model and its Application to Unsupervised Named-Entity Recognition

Angli Liu, Jingfei Du, Veselin Stoyanov

专题命中 其他LLM :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments NAACL 2019; updated to cite Zhou et al. (2018) EMNLP as a piece of related work

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.04444 2018-12-11 cs.CL cs.AI cs.LG stat.ML 82%

Character-Level Language Modeling with Deeper Self-Attention

Rami Al-Rfou, Dokook Choe, Noah Constant, Mandy Guo, Llion Jones

专题命中 其他LLM :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.02306 2018-09-10 cs.CL cs.AI cs.LG 82%

Unsupervised Cross-lingual Word Embedding by Multilingual Neural Language Models

Takashi Wada, Tomoharu Iwata

专题命中 其他LLM :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1507.01193 2015-07-07 cs.CL cs.AI cs.LG 82%

Dependency Recurrent Neural Language Models for Sentence Completion

Piotr Mirowski, Andreas Vlachos

专题命中 其他LLM :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted for publication at ACL 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10061 2026-05-12 cs.CL cs.AI 82%

Not-So-Strange Love: Language Models and Generative Linguistic Theories are More Compatible than They Appear

并非奇怪的爱:语言模型和生成语言理论比看起来更兼容

R. Thomas McCoy

专题命中 其他LLM :language model(title,abstract);分类 cs.CL、cs.AI

AI总结 本文探讨语言模型与生成语言理论的兼容性,提出LMs可支持基于形式结构的理论,扩展了可测试理论的范围,促进使用基础与生成理论的融合。

Comments Accepted to Behavioral and Brain Sciences; 4 pages; Commentary on "How Linguistics Learned to Stop Worrying and Love the Language Models" by Richard Futrell and Kyle Mahowald

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01693 2025-11-11 cs.LG cs.AI 82%

GPT, But Backwards: Exactly Inverting Language Model Outputs

Adrians Skapars, Edoardo Manino, Youcheng Sun, Lucas C. Cordeiro

专题命中 其他LLM :language model(title,abstract);分类 cs.AI、cs.LG;foundation model(comments)

Comments 7 pages, ICML 2025 Workshop on Reliable and Responsible Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16365 2024-12-24 cs.CL cs.AI 82%

Overview of the First Workshop on Language Models for Low-Resource Languages (LoResLM 2025)

Hansi Hettiarachchi, Tharindu Ranasinghe, Paul Rayson, Ruslan Mitkov, Mohamed Gaber, Damith Premasiri, Fiona Anting Tan, Lasitha Uyangodage

专题命中 其他LLM :language model(title,abstract);分类 cs.CL、cs.AI

Comments The First Workshop on Language Models for Low-Resource Languages (LoResLM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.15772 2024-07-19 cs.LG cs.AI cs.CY 82%

Ecosystem Graphs: The Social Footprint of Foundation Models

Rishi Bommasani, Dilara Soylu, Thomas I. Liao, Kathleen A. Creel, Percy Liang

专题命中 其他LLM :foundation model(title,abstract);分类 cs.AI、cs.LG

Comments Authored by the Center for Research on Foundation Models (CRFM) at the Stanford Institute for Human-Centered Artificial Intelligence (HAI). Ecosystem Graphs available at https://crfm.stanford.edu/ecosystem-graphs/

Journal ref Published in AIES 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.08975 2021-10-26 cs.CL cs.AI 82%

Deep Transfer Learning & Beyond: Transformer Language Models in Information Systems Research

Ross Gruetzemacher, David Paradice

专题命中 其他LLM :language model(title,abstract);分类 cs.CL、cs.AI

Comments Under review (revised once). Section 2, the literature review on deep transfer learning and transformer language models, is a valuable introduction for a broad audience (not just information systems researchers). 33 pages plus 13-page appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22708 2026-08-25 cs.AI 新提交 81%

CacheRouter: A Dual-Path Tool Routing Architecture with Cache-Preserving Main-Model Isolation for Long-Tail Tool Discovery

CacheRouter:一种用于长尾工具发现的保留缓存的主模型隔离双路径工具路由架构

Donghui Zha, Lingwei Xu, Linxiao Wu, Yixue Dong, Haochen Li

机构 * School of Mathematical Sciences, Beijing University of Posts and Telecommunications(北京邮电大学数学科学学院) School of Mathematics and Statistics, Chongqing University(重庆大学数学与统计学院) School of Science, Beijing Forestry University(北京林业大学理学院)

专题命中 其他LLM :LLM(summary_cn,abstract);分类 cs.AI

AI总结 针对LLM工具使用中渐进式披露与提示词缓存的权衡,提出CacheRouter双路径路由架构,通过主模型隔离与工具路由通道设计提升缓存命中率,降低输入成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07356 2026-08-25 cs.CL 版本更新 81%

Safety-Aligned Weights Are Not Enough: Refusal-Teacher-Guided Finetuning Enhances Safety and Downstream Performance under Harmful Finetuning Attacks

仅安全对齐权重不够:拒绝教师引导微调可在有害微调攻击下提升安全性与下游性能

Seokil Ham, Yubin Choi, Yujin Yang, Seungju Cho, Younghun Kim, Changick Kim

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

专题命中 其他LLM :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 针对有害微调攻击威胁,该研究提出拒绝教师引导微调框架,将安全FaaS微调范式从安全对齐权重微调转为基础权重在安全教师指导下的微调,可减少有害输出并提升下游任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20637 2026-08-24 cs.CR cs.AI 新提交 81%

ARQ: Agentic CodeQL Query Refinement for C/C++ Vulnerability Detection

ARQ:用于C/C++漏洞检测的智能CodeQL查询优化框架

Chunyi Wang, Yunfei Ke, Junfeng Yang, Yun-Yun Tsai, Penghui Li

专题命中 其他LLM :LLM(summary_cn,abstract);分类 cs.AI

AI总结 本文提出ARQ智能框架,利用合成C/C++程序的执行证据和LLM优化CodeQL查询,无需标注数据等,优化后查询的真阳性最多提升119.8%,还修复了长期悬而未决的问题并发现新漏洞。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19527 2026-08-21 cs.HC cs.CL cs.SD 新提交 81%

Does Listening Matter? Backchanneling and Nodding in AI Clone

倾听重要吗?AI克隆中的反馈式回应与点头动作

Koji Inoue, Kazushi Kato, Tatsuya Kawahara, Shunichi Kasahara

专题命中 其他LLM :LLM(summary_cn,abstract);分类 cs.CL

AI总结 本研究将口头反馈式回应与头部点头整合到配备语音克隆及LLM响应的AI克隆中,经35人被试内研究证实,添加互动倾听行为可提升AI克隆的感知专注度、真实感与共同在场感,其保真度需涵盖倾听行为。

Comments This paper has been accepted to the Late-Breaking Results (LBR) track of the 28th International Conference on Multimodal Interaction (ICMI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11256 2026-08-21 cs.LG cs.CY 版本更新 81%

Why AI Detection Fails for Academic Integrity

为什么AI检测无法保障学术诚信

Jonathan A. Karr, Grigorii Khvatskii, Ting Hua, Nitesh V. Chawla

机构 * University of Notre Dame(圣母大学)

专题命中 其他LLM :LLM(summary_cn,abstract);分类 cs.LG

AI总结 该研究发现商业AI检测器无法区分AI编辑与LLM完整生成稿,轻度AI辅助润色的论文标记率远高于未修改稿,且人性化处理可大幅规避检测,表明检测器分数不能单独作为学术不端证据。

Comments Accepted to ACM AI Leadership Summit

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18846 2026-08-20 cs.AI 新提交 81%

ORBITER: Conflict-Aware Decision-Making for Agentic Last-Mile Delivery

ORBITER:面向智能体的最后一公里配送的冲突感知决策方法

Mingzhao Li, Chenxi Liu, Yan Zhao, Hao Miao

专题命中 其他LLM :LLM(summary_cn,abstract);分类 cs.AI

AI总结 本文针对最后一公里配送决策的可解释性与可靠性问题,提出ORBITER框架,结合LLM与结构化决策机制,在四城市数据上较最优基线平均提升9.2%,验证了方法的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18545 2026-08-20 cs.CL 新提交 81%

Shared Circuits for Shared Grammar: Tracing Subject-Verb Agreement Across Languages

共享语法的共享电路:跨语言追踪主谓一致

Isabella Gidi, Antonio Almudévar, Core Francisco Park, Naomi Saphra, Ricard Marxer

机构 * Harvard University(哈佛大学) University of Zaragoza(萨拉戈萨大学) Boston University(波士顿大学) Univ Toulon(土伦大学) Aix Marseille Univ(艾克斯-马赛大学) CNRS(法国国家科学研究中心) LIS(信息科学实验室) ILLS(语言与语言科学研究所)

专题命中 其他LLM :large language model(abstract,abstract_cn);language model(abstract,abstract_cn);分类 cs.CL

AI总结 本文研究多语言大语言模型的跨语言共享机制,以主谓一致为对象,通过29种语言的实验发现其复用部分共享计算结构,且与屈折语言的电路更相似。

Comments 25 pages including appendices, 16 figures. Accepted to COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16934 2026-08-20 cs.AR cs.CL 版本更新 81%

SeqFeed: Improving Agentic RTL Code Generation with Sequential Behavior Feedback

SeqFeed:通过顺序行为反馈改进智能体RTL代码生成

Yuxin Du, Juxin Niu, Tao Hu, Xi Wang, Zhe Jiang, Nan Guan

专题命中 其他LLM :LLM(summary_cn,abstract);分类 cs.CL

AI总结 针对智能体RTL代码生成的时序信息传递难题,本研究提出含SeQuery和SeGraph机制的SeqFeed,通过满足三项反馈要求提升LLM的RTL代码生成通过率。

详情

展开后加载摘要…

URL PDF HTML 收藏