arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 7608 信号源:cs.CL, cs.AI, cs.LG

1. 知识编辑与模型理解 7608 篇

2605.09023 2026-05-12 cs.SE 82%

Using Semantic Distance to Estimate Uncertainty in LLM-Based Code Generation

利用语义距离估计基于大语言模型的代码生成的不确定性

Weilin He, Arindam Sharma, Cristina David

专题命中 知识编辑与模型理解 :LLM(title,abstract)

AI总结 本文提出基于语义距离的不确定性估计方法,通过衡量生成程序执行行为的差异性,提升代码生成的正确性评估效果,在多个基准测试中优于现有方法。

Comments Preprint. 23 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22040 2026-05-08 cs.CL cs.AI cs.LG 82%

Leviathan: Decoupling Input and Output Representations in Language Models

Leviathan:解耦语言模型的输入和输出表示

Reza T. Batley, Sourav Saha

机构 * Department of Aerospace and Ocean Engineering(航空航天与海洋工程系)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 Leviathan通过引入学习嵌入向量化方法,解耦语言模型的输入和输出表示,从而在参数增加极小的情况下提升语言建模性能,尤其在稀有词上表现显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27413 2026-04-27 cs.LG cs.AI cs.CL 82%

Atlas-Alignment: Making Interpretability Transferable Across Language Models

Atlas-Alignment:使语言模型间的可解释性可迁移

Bruno Puri, Jim Berend, Sebastian Lapuschkin, Wojciech Samek

机构 * Department of Artificial Intelligence, Fraunhofer Heinrich Hertz Institute(人工智能系,弗劳恩霍夫 Heinrich Hertz 研究所) Department of Electrical Engineering and Computer Science, Technische Universität Berlin(电气工程与计算机科学系,柏林技术大学) Centre of eXplainable Artificial Intelligence, Technological University Dublin(可解释人工智能中心,都柏林技术大学) BIFOLD - Berlin Institute for the Foundations of Learning and Data(BIFOLD - 柏林学习与数据基础研究所)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 Atlas-Alignment通过在新模型的潜在空间与预训练的Concept Atlas对齐,实现无需标注数据的可解释性迁移,降低可解释AI的成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04972 2026-04-08 cs.CV 82%

RCP: Representation Consistency Pruner for Mitigating Distribution Shift in Large Vision-Language Models

RCP:一种用于缓解大视觉-语言模型中分布偏移的表示一致性剪枝方法

Jianwei Zhang, Chaoning Zhang, Sihan Cao, Wang Liu, Pengcheng Zheng, Jiaxin Huang, Caiyan Qin, Yalan Ye, Wei Dong, Yang Yang

机构 * School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院) Machine Learning Department, Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学机器学习系) School of Robotics and Advanced Manufacture, Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)机器人与先进制造学院)

专题命中 知识编辑与模型理解 :language model(title,abstract);LLM(abstract)

AI总结 本文提出RCP框架,通过累积视觉token剪枝与延迟修复机制,有效缓解大视觉-语言模型中的分布偏移问题,实验表明可剪除高达88.9%的视觉token并减少85.7%的FLOPs,同时保持较低的精度损失。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23095 2026-04-07 cs.CV 82%

Revisiting Multimodal Positional Encoding in Vision-Language Models

重新审视视觉-语言模型中的多模态位置编码

Jie Huang, Xuejing Liu, Sibo Song, Ruibing Hou, Hong Chang, Junyang Lin, Shuai Bai

机构 * Qwen Team, Alibaba Group(Qwen团队,阿里巴巴集团) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

专题命中 知识编辑与模型理解 :language model(title,abstract);LLM(abstract)

AI总结 本文分析了多模态Rotary Positional Embedding的核心组件,提出MHRoPE和MRoPE-Interleave两种改进方法,在多种基准测试中表现优异,提升了多模态理解能力。

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02459 2026-04-06 cs.LG cs.AI cs.CL 82%

On the Geometric Structure of Layer Updates in Deep Language Models

深度语言模型中层更新的几何结构研究

Jun-Sik Yoo

机构 * Institute of Basic Science, Korea University(韩国大学基础科学研究所)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文研究深度语言模型中层更新的几何结构,发现层更新可分解为主导的tokenwise组件和一个几何上不同的残差部分,残差对输出扰动有显著影响。

Comments 11 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23821 2026-03-26 cs.CL cs.AI cs.LG 82%

Perturbation: A simple and efficient adversarial tracer for representation learning in language models

扰动:一种简单且高效的对抗追踪器,用于语言模型中的表示学习

Joshua Rozner, Cory Shain

机构 * Stanford University(斯坦福大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出扰动方法,通过微调语言模型于单个对抗示例,揭示训练模型在不同语言粒度上的结构迁移,展示语言模型在无监督学习中获得语言抽象的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23037 2026-03-25 cs.CV cs.AI cs.CL cs.LG cs.RO 82%

YOLOv10 with Kolmogorov-Arnold networks and vision-language foundation models for interpretable object detection and trustworthy multimodal AI in computer vision perception

YOLOv10结合Kolmogorov-Arnold网络和视觉-语言基础模型用于可解释的目标检测和可信的多模态AI在计算机视觉感知

Marios Impraimakis, Daniel Vazquez, Feiyu Zhou

机构 * University of Bath(巴斯大学) Zhejiang University(浙江大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出一种基于Kolmogorov-Arnold网络和视觉-语言基础模型的YOLOv10框架,通过七种几何和语义特征提升目标检测的可解释性,并在模糊、遮挡或低纹理场景中实现可信的置信度估计。

Comments 14 pages, 23 Figures, 6 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00490 2026-03-17 cs.NE 82%

LLM-Driven Instance-Specific Heuristic Generation and Selection

基于大语言模型的实例特定启发式生成与选择

Shaofeng Zhang, Shengcai Liu, Ning Lu, Jiahao Wu, Ji Liu, Yew-Soon Ong, Ke Tang

专题命中 知识编辑与模型理解 :LLM(title);large language model(abstract);language model(abstract)

AI总结 本文提出InstSpecHH框架,通过实例特征划分问题子类,实现差异化启发式设计,提升求解效率并减少计算开销。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10071 2026-03-12 cs.LG cs.AI cs.CL 82%

Dissecting Chronos: Sparse Autoencoders Reveal Causal Feature Hierarchies in Time Series Foundation Models

拆解Chronos:稀疏自编码器揭示时间序列基础模型中的因果特征层次

Anurag Mishra

机构 * Rochester Institute of Technology(罗切斯特理工学院)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文通过稀疏自编码器揭示时间序列基础模型中的因果特征层次,发现中编码器的特征对预测质量影响最大,表明Chronos-T5依赖突变动态而非周期性模式识别。

Comments Accepted as a poster in ICLR 2026 Workshop on Time Series in the Age of Large Models (TSALM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06081 2026-03-09 cs.CV 82%

Lyapunov Probes for Hallucination Detection in Large Foundation Models

Lyapunov探针用于大型基础模型中的幻觉检测

Bozhi Luan, Gen Li, Yalan Qin, Jifeng Guo, Yun Zhou, Faguo Wu, Hongwei Zheng, Wenjun Wu, Zhaoxin Fan

机构 * Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, School of Artificial Intelligence, Beihang University(北京未来区块链与隐私计算先进创新中心,人工智能学院,北航) School of Electronic and Information Engineering, State Key Laboratory of CNS/ATM, Beihang University(电子与信息工程学院, CNS/ATM 国家重点实验室,北航) National Key Laboratory of Information Systems Engineering, National University of Defense Technology(信息系统工程国家重点实验室,国防科技大学) Beijing Academy of Blockchain and Edge Computing(北京区块链与边缘计算研究院) Shanghai University(上海大学)

专题命中 知识编辑与模型理解 :foundation model(title);large language model(abstract);language model(abstract)

AI总结 通过Lyapunov探针检测大型基础模型中的幻觉,利用动力系统稳定性理论分析知识过渡区域的边界特征。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00102 2026-03-03 cs.RO 82%

Designing Social Robots with Ethical, User-Adaptive Explainability in the Era of Foundation Models

在基础模型时代设计具有伦理性和用户适应性的社交机器人

Fethiye Irmak Dogan, Alva Markelius, Hatice Gunes

机构 * University of Cambridge(剑桥大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);LLM(abstract)

AI总结 本文提出在基础模型驱动的社交机器人中,需通过伦理性和用户适应性的可解释性设计,解决适应与解释委托给基础模型带来的挑战。

Comments Companion Proceedings of the 21st ACM/IEEE International Conference on Human-Robot Interaction

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14896 2026-03-02 cs.CV 82%

Leveraging Multimodal LLM Descriptions of Activity for Explainable Semi-Supervised Video Anomaly Detection

利用多模态大语言模型对活动的描述进行可解释的半监督视频异常检测

Furkan Mumcu, Michael J. Jones, Anoop Cherian, Yasin Yilmaz

机构 * Department of Electrical Engineering University of South Florida(电气工程系 佛罗里达州立大学) Mitsubishi Electric Research Laboratories (MERL)(三菱电机研究实验室(MERL))

专题命中 知识编辑与模型理解 :LLM(title);large language model(abstract);language model(abstract)

AI总结 本文提出利用多模态大语言模型描述活动,以实现可解释的半监督视频异常检测,有效检测复杂交互异常并提升模型可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11246 2026-02-13 cs.LG cs.AI cs.CL cs.IT math.CO math.IT 82%

How Many Features Can a Language Model Store Under the Linear Representation Hypothesis?

在线性表示假设下,语言模型能存储多少特征?

Nikhil Garg, Jon Kleinberg, Kenny Peng

机构 * Cornell University(康奈尔大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 该研究在线性表示假设下,探讨了语言模型存储特征的理论上限和下限,证明了神经元可存储指数级特征数量,并揭示了线性可访问性比线性表示更强的假设。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08579 2026-02-10 cs.CL cs.AI cs.LG 82%

Training Language Models to Explain Their Own Computations

训练语言模型以解释其自身的计算

Belinda Z. Li, Zifan Carl Guo, Vincent Huang, Jacob Steinhardt, Jacob Andreas

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究通过训练语言模型解释自身计算,发现其能有效生成内部机制的自然语言描述,为可解释性方法提供新途径。

Comments 23 pages, 8 tables, 7 figures. Code and data at https://github.com/TransluceAI/introspective-interp

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14622 2026-01-22 cs.RO 82%

Probing Prompt Design for Socially Compliant Robot Navigation with Vision Language Models

通过视觉语言模型探索社交合规机器人导航的提示设计

Ling Xiao, Toshihiko Yamasaki

机构 * Hokkaido University(北海道大学) The University of Tokyo(东京大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract)

AI总结 本文通过设计社交合规的提示策略,提升小型视觉语言模型在机器人导航中的行动准确性,发现与自我竞争的提示设计效果最佳。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06616 2026-01-13 cs.HC 82%

LLM-Driven Accessible Interface: A Model-Based Approach

基于大语言模型的可及性界面:一种基于模型的方法

Blessing Jerry, Lourdes Moreno, Virginia Francisco, Raquel Hervas

专题命中 知识编辑与模型理解 :LLM(title);large language model(abstract);language model(abstract)

AI总结 本文提出了一种基于模型的框架,利用大语言模型生成个性化、多模态且符合可及性标准的用户界面,通过结构化用户档案、声明性适应规则和验证提示模板,提升医疗场景中对认知和感官需求的可及性支持。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02072 2025-12-03 hep-ph physics.data-an 82%

QCD in Language Models: What do they really know about QCD?

语言模型中的QCD:它们真的了解QCD吗?

Antonin Sulc, Patrick L. S. Connor

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract)

AI总结 研究分析了语言模型对QCD的理解,揭示其在参数空间中对QCD概念的嵌入模式,并指出模型在高级量子场论表示上的局限性。

Comments 6 pages, 4 figures, presented at EPS HEP 2025 by Patrick L.S. Connor as Oral Presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11770 2025-11-12 cs.LG cs.AI cs.CL stat.ML 82%

Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors

Jing Huang, Junyi Tao, Thomas Icard, Diyi Yang, Christopher Potts

机构 * stanford(斯坦福大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23184 2025-10-28 cs.CV 82%

Finding 3D Scene Analogies with Multimodal Foundation Models

Junho Kim, Young Min Kim

机构 * Institute of New Media and Communications(新媒体与通讯研究所) Dept. of Electrical and Computer Engineering(电气与计算机工程系)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);language model(abstract)

Comments Accepted to FM4RoboPlan workshop at RSS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17349 2025-10-02 cs.CV 82%

Beyond Semantics: Rediscovering Spatial Awareness in Vision-Language Models

Jianing Qi, Jiawei Liu, Hao Tang, Zhigang Zhu

机构 * CUNY Graduate Center(纽约大学研究生中心) Borough of Manhattan Community College(曼哈顿社区学院) The City College of New York(纽约城市学院)

专题命中 知识编辑与模型理解 :language model(title,abstract);LLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05230 2025-09-11 cs.CL cs.AI cs.LG 82%

CURE: Controlled Unlearning for Robust Embeddings -- Mitigating Conceptual Shortcuts in Pre-Trained Language Models

Aysenur Kocak, Shuo Yang, Bardh Prenkaj, Gjergji Kasneci

机构 * Technical University of Munich(慕尼黑技术大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at the Conference on Empirical Methods in Natural Language Processing (EMNLP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14729 2025-08-01 cs.CV 82%

Uncovering Cultural Representation Disparities in Vision-Language Models

Ram Mohan Rao Kadiyala, Siddhant Gupta, Jebish Purbey, Srishti Yadav, Suman Debnath, Alejandro Salamanca, Desmond Elliott

机构 * Cohere Labs Community IIT Roorkee(印度理工学院罗尔克hee分校) University of Copenhagen(哥本哈根大学) Cohere Labs Amazon Datasets(亚马逊数据集)

专题命中 知识编辑与模型理解 :language model(title,abstract);prompting(abstract)

Comments 28 pages, 36 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15836 2025-07-30 cs.CV 82%

VLM-CPL: Consensus Pseudo Labels from Vision-Language Models for Annotation-Free Pathological Image Classification

Lanfeng Zhong, Zongyao Huang, Yang Liu, Wenjun Liao, Shichuan Zhang, Guotai Wang, Shaoting Zhang

机构 * School of Mechanical and Electrical Engineering, University of Electronic Science and Technology of China(电子科技大学机械与电子工程学院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Department of Pathology, Sichuan Clinical Research Center for Cancer, Sichuan Cancer Hospital & Institute, Affiliated Cancer Hospital of University of Electronic Science and Technology of China(pathology department, 四川省癌症临床研究中心, 四川省肿瘤医院及研究所, 电子科技大学附属肿瘤医院) Department of Radiation Oncology, Sichuan Cancer Hospital and Institute, University of Electronic Science and Technology of China(放射肿瘤科, 四川省肿瘤医院及研究所, 电子科技大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);prompting(abstract)

Comments Accepted at TMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18631 2025-07-28 cs.CR 82%

Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment

Hao Li, Lijun Li, Zhenghao Lu, Xianyi Wei, Rui Li, Jing Shao, Lei Sha

专题命中 知识编辑与模型理解 :LLM(title,abstract);post-training(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09081 2025-07-15 cs.CV 82%

From Physics to Foundation Models: A Review of AI-Driven Quantitative Remote Sensing Inversion

Zhenyu Yu, Mohd Yamani Idna Idris, Hua Wang, Pei Wang, Junyi Chen, Kun Wang

专题命中 知识编辑与模型理解 :foundation model(title,abstract);pretraining(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07486 2025-07-11 q-bio.OT 82%

Sparse Autoencoders Reveal Interpretable Structure in Small Gene Language Models

Haoxiang Guan, Jiyan He, Jie Zhang

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract)

Comments AI4X 2025 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01790 2025-07-03 cs.CL cs.AI cs.CV cs.LG 82%

How Do Vision-Language Models Process Conflicting Information Across Modalities?

Tianze Hua, Tian Yun, Ellie Pavlick

机构 * Brown University(布朗大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments All code and resources are available at: https://github.com/ethahtz/vlm_conflicting_info_processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20178 2025-06-26 cs.CL cs.AI cs.LG 82%

COIN: Uncertainty-Guarding Selective Question Answering for Foundation Models with Provable Risk Guarantees

Zhiyuan Wang, Jinhao Duan, Qingni Wang, Xiaofeng Zhu, Tianlong Chen, Xiaoshuang Shi, Kaidi Xu

机构 * University of Electronic Science and Technology of China(电子科技大学) Drexel University(德雷塞尔大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01755 2025-06-18 cs.CY 82%

AI Language Models Could Both Help and Harm Equity in Marine Policymaking: The Case Study of the BBNJ Question-Answering Bot

Matt Ziegler, Sarah Lothian, Brian O'Neill, Richard Anderson, Yoshitaka Ota

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract)

Journal ref npj Ocean Sustainability | (2025)4:32

详情

展开后加载摘要…

URL PDF HTML 收藏