arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 7608 信号源:cs.CL, cs.AI, cs.LG

1. 知识编辑与模型理解 7608 篇

2405.12522 2024-05-22 cs.CL cs.LG 84%

Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models

Charles O'Neill, Thang Bui

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08744 2024-05-07 cs.CL cs.LG 84%

Circuit Component Reuse Across Tasks in Transformer Language Models

Jack Merullo, Carsten Eickhoff, Ellie Pavlick

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.LG

Comments Accepted at ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12145 2024-04-19 cs.CL cs.AI 84%

From Form(s) to Meaning: Probing the Semantic Depths of Language Models Using Multisense Consistency

Xenia Ohmer, Elia Bruni, Dieuwke Hupkes

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.06948 2024-04-12 cs.CL cs.AI 84%

MetaCheckGPT -- A Multi-task Hallucination Detector Using LLM Uncertainty and Meta-models

Rahul Mehta, Andrew Hoblitzell, Jack O'Keefe, Hyeju Jang, Vasudeva Varma

专题命中 知识编辑与模型理解 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments Entry for SemEval-2024 Shared Task 6: SHROOM, a Shared-task on Hallucinations and Related Observable Overgeneration Mistakes

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03044 2024-04-05 cs.LG cs.AI 84%

The Artificial Intelligence Ontology: LLM-assisted construction of AI concept hierarchies

Marcin P. Joachimiak, Mark A. Miller, J. Harry Caufield, Ryan Ly, Nomi L. Harris, Andrew Tritt, Christopher J. Mungall, Kristofer E. Bouchard

专题命中 知识编辑与模型理解 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.02060 2024-03-19 cs.CL cs.LG 84%

Representation Deficiency in Masked Language Modeling

Yu Meng, Jitin Krishnan, Sinong Wang, Qifan Wang, Yuning Mao, Han Fang, Marjan Ghazvininejad, Jiawei Han, Luke Zettlemoyer

专题命中 知识编辑与模型理解 :language model(title,abstract);pretraining(abstract);分类 cs.CL、cs.LG

Comments ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03028 2024-03-06 cs.AI cs.CL 84%

Word Importance Explains How Prompts Affect Language Model Outputs

Stefan Hackmann, Haniyeh Mahmoudian, Mark Steadman, Michael Schmidt

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.06991 2024-02-06 cs.LG cs.CL stat.ML 84%

Unsupervised Contrast-Consistent Ranking with Language Models

Niklas Stoehr, Pengxiang Cheng, Jing Wang, Daniel Preotiuc-Pietro, Rajarshi Bhowmik

专题命中 知识编辑与模型理解 :language model(title,abstract);prompting(abstract);分类 cs.CL、cs.LG

Comments Long Paper at EACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08027 2023-12-14 cs.CL cs.AI 84%

Helping Language Models Learn More: Multi-dimensional Task Prompt for Few-shot Tuning

Jinta Weng, Jiarui Zhang, Yue Hu, Daidong Fa, Xiaofeng Xuand, Heyan Huang

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: text overlap with arXiv:2210.16489

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.10905 2023-11-21 cs.CL cs.AI 84%

Flexible Model Interpretability through Natural Language Model Editing

Karel D'Oosterlinck, Thomas Demeester, Chris Develder, Christopher Potts

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

Comments Extended Abstract -- work in progress. BlackboxNLP2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09053 2023-11-16 cs.CL cs.AI 84%

Assessing Knowledge Editing in Language Models via Relation Perspective

Yifan Wei, Xiaoyan Yu, Huanhuan Ma, Fangyu Lei, Yixuan Weng, Ran Song, Kang Liu

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.07383 2023-11-14 cs.CL cs.LG 84%

LM-Polygraph: Uncertainty Estimation for Language Models

Ekaterina Fadeeva, Roman Vashurin, Akim Tsvigun, Artem Vazhentsev, Sergey Petrakov, Kirill Fedyanin, Daniil Vasilev, Elizaveta Goncharova, Alexander Panchenko, Maxim Panov, Timothy Baldwin, Artem Shelmanov

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.LG

Comments Accepted at EMNLP-2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.15910 2023-10-25 cs.CL cs.AI 84%

Characterizing Mechanisms for Factual Recall in Language Models

Qinan Yu, Jack Merullo, Ellie Pavlick

专题命中 知识编辑与模型理解 :language model(title,abstract);pretraining(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.14815 2023-09-19 cs.CL cs.LG 84%

Black-box language model explanation by context length probing

Ondřej Cífka, Antoine Liutkus

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.LG

Comments 11 pages, 9 figures. ACL 2023 short paper camera-ready. Demos at https://cifkao.github.io/context-probing/ and https://huggingface.co/spaces/cifkao/context-probing ; code at https://github.com/cifkao/context-probing/

Journal ref Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) (2023), 1067--1079

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.15430 2023-03-30 cs.CL cs.LG 84%

TextMI: Textualize Multimodal Information for Integrating Non-verbal Cues in Pre-trained Language Models

Md Kamrul Hasan, Md Saiful Islam, Sangwu Lee, Wasifur Rahman, Iftekhar Naim, Mohammed Ibrahim Khan, Ehsan Hoque

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06376 2022-10-13 cs.CL cs.AI 84%

Probing Commonsense Knowledge in Pre-trained Language Models with Sense-level Precision and Expanded Vocabulary

Daniel Loureiro, Alípio Mário Jorge

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.12153 2022-09-27 cs.CL cs.AI 84%

WinoDict: Probing language models for in-context word acquisition

Julian Martin Eisenschlos, Jeremy R. Cole, Fangyu Liu, William W. Cohen

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.10744 2022-06-23 cs.CL cs.AI 84%

Don't Forget About Pronouns: Removing Gender Bias in Language Models Without Losing Factual Gender Information

Tomasz Limisiewicz, David Mareček

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

Comments Presented at GeBNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.10198 2022-04-25 cs.CL cs.AI 84%

Context-Aware Language Modeling for Goal-Oriented Dialogue Systems

Charlie Snell, Mengjiao Yang, Justin Fu, Yi Su, Sergey Levine

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.03489 2022-04-08 cs.CL cs.LG 84%

Position-based Prompting for Health Outcome Generation

M. Abaho, D. Bollegala, P. Williamson, S. Dodd

专题命中 知识编辑与模型理解 :prompting(title,abstract);language model(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.02647 2022-02-08 cs.CL cs.AI 84%

Ethics, Rules of Engagement, and AI: Neural Narrative Mapping Using Large Transformer Language Models

Philip Feldman, Aaron Dant, David Rosenbluth

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

Comments 18 Pages, 13 figures

Journal ref Bulletin of the Technical Committee on Data Engineering, Vol. 44 No. 4 December 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.00781 2020-07-22 cs.CL cs.LG stat.ML 84%

On the comparability of Pre-trained Language Models

Matthias Aßenmacher, Christian Heumann

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.LG

Journal ref Proceedings of the 5th Swiss Text Analytics Conference (SwissText) & 16th Conference on Natural Language Processing (KONVENS), Zurich, Switzerland, June 23-25, 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08744 2024-09-16 cs.CV cs.LG 84%

Uncertainty and Generalizability in Foundation Models for Earth Observation

Raul Ramos-Pollan, Freddie Kalaitzis, Karthick Panner Selvam

专题命中 知识编辑与模型理解 :foundation model(title,abstract);pretraining(abstract);分类 cs.LG

Comments A large ablation study measuring uncertainty and spatial generalizability with 8 foundation models, 11 world regions and 7 downstream tasks

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13900 2026-08-17 cs.DB cs.AI cs.CL cs.LG 新提交 83%

Agentic Transaction: Towards ACID-Compliant Agent Systems

智能体事务:面向ACID兼容的智能体系统

Zhaoyan Sun, Xiaoxiao Wang, Guoliang Li

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 该研究提出ACID兼容的智能体事务框架,开发对应数据智能体,在基准测试中较含Claude Code的现有智能体提升10.6%,为构建可信可扩展AI智能体开辟新方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16616 2026-08-13 cs.DB 版本更新 83%

LDI: Localized Data Imputation for Text-Rich Tables

LDI:面向文本丰富表的局部数据填补

Soroush Omidvartehrani, Davood Rafiei

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract_cn);large language model(abstract);language model(abstract)

AI总结 本文提出LDI框架,通过局部推理利用LLM填补文本丰富表中的缺失值,提升准确性和可解释性,实验证明其优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04765 2026-05-29 cs.CL cs.AI cs.LG physics.comp-ph 83%

Differential syntactic and semantic encoding in LLMs

大型语言模型中句法与语义的差异编码

Santiago Acevedo, Alessandro Laio, Marco Baroni

机构 * Catalan Institute of Research and Advanced Studies (ICREA) and Universitat Pompeu Fabra (UPF)(加泰罗尼亚研究与高级科学研究所(ICREA)和庞培法华大学(UPF))

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究通过平均共享句法结构或语义的句子隐藏表示向量,发现大型语言模型(以DeepSeek-V3为例)的内部层表示中句法和语义信息至少部分线性编码,且两者编码轮廓不同,可一定程度解耦。

Comments Published as conference paper at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24946 2026-05-26 cs.CV 83%

Interpretability Transfer from Language to Vision via Sparse Autoencoders

通过稀疏自编码器实现从语言到视觉的可解释性迁移

Alexey Kravets, Da Li, Chuan Li, Da Chen, Vinay P. Namboodiri

机构 * University of Bath, UK(巴斯大学) Lambda, Inc.(Lambda公司) Samsung AI Centre Cambridge(三星AI研究中心)

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);language model(abstract)

AI总结 提出VISTA框架,通过约束视觉投影器将视觉token映射到LLM的文本SAE空间,实现无需专用视觉SAE的视觉可解释性,并在对象移除和替换任务上分别提升35%和47%。

Journal ref ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23765 2026-05-08 cs.CL cs.AI cs.LG 83%

Knowledge-Level Consistency Reinforcement Learning: Dual-Fact Alignment for Long-Form Factuality

知识级一致性强化学习:长形式事实性的双事实对齐

Junliang Li, Yucheng Wang, Yan Chen, Yu Ran, Ruiqing Zhang, Jing Liu, Hua Wu, Haifeng Wang

机构 * Baidu Inc.(百度公司)

专题命中 知识编辑与模型理解 :RLHF(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出KLCF框架,通过双事实对齐机制提升长形式生成的事实性,有效缓解幻觉和保守倾向,提升精度与召回率。

Comments 32 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22779 2026-04-28 cs.LG cs.AI cs.CL 83%

KARL: Mitigating Hallucinations in LLMs via Knowledge-Boundary-Aware Reinforcement Learning

KARL:通过知识边界感知强化学习缓解大语言模型的幻觉

Cheng Gao, Cheng Huang, Kangyang Luo, Ziqing Qiao, Shuzheng Si, Huimin Chen, Chaojun Xiao, Maosong Sun

机构 * Tsinghua University(清华大学)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 KARL通过知识边界感知强化学习框架,动态调整大语言模型的回避行为,有效抑制幻觉并保持高精度。

Comments 21 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06032 2026-01-13 cs.HC 83%

Applied Theory of Mind and Large Language Models -- how good is ChatGPT at solving social vignettes?

应用认知理论与大语言模型 -- ChatGPT在解决社会情境题方面表现如何?

Anna Katharina Holl-Etten, Nina Schnaderbeck, Elizaveta Kosareva, Leonhard Aron Prattke, Ralph Krueger, Lisa Marie Warner, Nora C. Vetter

专题命中 知识编辑与模型理解 :large language model(title);language model(title)

AI总结 研究评估了GPT-4在解决社会情境题方面的表现,发现其在高阶认知理论任务中接近人类水平,但在不确定性标记使用上仍需进一步优化。

Comments 40 pages, 6 figures, 3 supplements

详情

展开后加载摘要…

URL PDF HTML 收藏