arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 7583 信号源:cs.CL, cs.AI, cs.LG

1. 知识编辑与模型理解 7583 篇

2512.15708 2025-12-18 cs.CV 78%

Multi-View Foundation Models

多视图基础模型

Leo Segre, Or Hirschorn, Shai Avidan

机构 * Tel Aviv University(特拉维夫大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

AI总结 本文提出一种将基础模型转换为多视图基础模型的方法,通过引入3D感知注意力层提升多视角特征一致性,应用于表面法线估计和多视角分割任务,实验表明其在特征匹配上有显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09321 2025-12-16 cs.CR 78%

ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data

ObliInjection: 面向多源数据LLM代理的顺序无关提示注入攻击

Reachal Wang, Yuqi Jia, Neil Zhenqiang Gong

专题命中 知识编辑与模型理解 :LLM(title,abstract)

AI总结 ObliInjection是一种针对多源数据LLM代理的新型提示注入攻击,通过顺序无关损失和顺序GCG算法有效污染输入数据以误导模型执行攻击者指定任务。

Comments To appear in NDSS 2026. For slides, see https://people.duke.edu/~zg70/code/PromptInjection.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11104 2025-12-15 cs.CV 78%

Information-driven Fusion of Pathology Foundation Models for Enhanced Disease Characterization

基于信息驱动的病理基础模型融合以增强疾病表征

Brennan Flannery, Thomas DeSilvio, Jane Nguyen, Satish E. Viswanath

机构 * Case Western Reserve University(凯斯西储大学) Department of Biomedical Engineering(生物医学工程系) Cleveland Clinic(克利夫兰诊所) Department of Pathology(病理学系) Emory University(埃默里大学) Department of Pediatrics(儿科学系) Louis Stokes VA Cleveland Medical Center(路易斯·斯托克斯退伍军人医疗中心)

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

AI总结 本研究提出基于信息驱动的病理基础模型融合方法,通过智能融合提升癌症分级和分期的预测性能与可解释性。

Comments 29 Pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08439 2025-12-10 cs.CV 78%

LapFM: A Laparoscopic Segmentation Foundation Model via Hierarchical Concept Evolving Pre-training

LapFM:通过分层概念演化的预训练构建腹腔镜分割基础模型

Qing Xu, Kun Yuan, Yuxiang Luo, Yuhao Zhai, Wenting Duan, Nassir Navab, Zhen Chen

机构 * School of Computer Science, University of Lincoln, UK(英国林肯大学计算机科学学院) University of Nottingham, UK(英国诺丁汉大学) University of Nottingham Ningbo China, China(中国宁波诺丁汉大学) University of Strasbourg, France(法国斯特拉斯堡大学) Technical University of Munich, Germany(德国慕尼黑技术大学) Graduate School of Information, Production and Systems, Waseda University, Japan(日本早稻田大学信息、生产与系统研究生院) Department of Gastrointestinal Surgery, The Second Qilu Hospital, Shandong University, China(中国山东大学第二齐鲁医院胃肠外科) School of Engineering and Physical Science, University of Lincoln, Lincoln LN6 7TS, UK(英国林肯大学工程与物理科学学院) Yale University, New Haven, CT 06510, USA(美国耶鲁大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

AI总结 LapFM通过分层概念演化预训练方法,构建了基于腹腔镜手术图像的大型基准,实现了对复杂手术场景的高效分割和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22019 2025-12-09 cs.CV 78%

Intra-Class Probabilistic Embeddings for Uncertainty Estimation in Vision-Language Models

类内概率嵌入用于视觉-语言模型中的不确定性估计

Zhenxiang Lin, Maryam Haghighat, Will Browne, Dimity Miller

机构 * Queensland University of Technology(昆士兰理工大学)

专题命中 知识编辑与模型理解 :language model(title,abstract)

AI总结 本研究提出一种无需训练的后处理方法,通过类内概率嵌入提升视觉-语言模型的不确定性估计,有效检测错误预测。

Comments Accepted at the IEEE/CVF Winter Conference on Applications of Computer Vision 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19532 2025-12-08 q-bio.BM 78%

Toward the Explainability of Protein Language Models

迈向蛋白质语言模型的可解释性

Andrea Hunklinger, Noelia Ferruz

专题命中 知识编辑与模型理解 :language model(title,abstract)

AI总结 本文探讨了XAI在蛋白质语言模型中的应用,提出了XAI在蛋白质研究中的五个潜在角色,并呼吁推动可解释性的发展。

Comments 15 pages, 6 figures; version 4: Additional revision of the manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17162 2025-12-08 cs.CR 78%

Analyzing PDFs like Binaries: Adversarially Robust PDF Malware Analysis via Intermediate Representation and Language Model

像二进制一样分析PDF:通过中间表示和语言模型实现对抗鲁棒的PDF恶意软件分析

Side Liu, Jiang Ming, Guodong Zhou, Xinyi Liu, Jianming Fu, Guojun Peng

专题命中 知识编辑与模型理解 :language model(title,abstract)

AI总结 通过中间表示和语言模型实现对抗鲁棒的PDF恶意软件分析,利用语义和结构特征提取提升检测性能。

Comments Accepted by ACM CCS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00493 2025-12-02 cs.CV 78%

CC-FMO: Camera-Conditioned Zero-Shot Single Image to 3D Scene Generation with Foundation Model Orchestration

CC-FMO:基于相机的零样本单图像到3D场景生成与基础模型协调

Boshi Tang, Henry Zheng, Rui Huang, Gao Huang

机构 * Tsinghua University(清华大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

AI总结 CC-FMO通过结合语义感知和结构化潜在表示,实现基于相机的零样本单图像到3D场景生成,提升场景连贯性和实例保真度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22171 2025-12-02 cs.CV 78%

HARMONY: Hidden Activation Representations and Model Output-Aware Uncertainty Estimation for Vision-Language Models

HARMONY:隐藏的激活表示和模型输出感知的不确定性估计用于视觉-语言模型

Erum Mushtaq, Zalan Fabian, Yavuz Faruk Bakman, Anil Ramakrishna, Mahdi Soltanolkotabi, Salman Avestimehr

机构 * University of Southern California(南加州大学) Amazon AGI(亚马逊人工智能实验室)

专题命中 知识编辑与模型理解 :language model(title,abstract)

AI总结 HARMONY通过整合生成token、模型输出不确定性分数和隐藏表示,提升视觉-语言模型的不确定性估计性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22961 2025-12-01 cs.CV 78%

HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model

HMR3D:用于大视觉-语言模型的层次多模态表示以实现3D场景理解

Chen Li, Eric Peh, Basura Fernando

机构 * Institute of High-Performance Computing, Agency for Science, Technology and Research, Singapore(高性能计算研究所,科技研究局,新加坡) Centre for Frontier AI Research, Agency for Science, Technology and Research, Singapore(前沿人工智能研究中心,科技研究局,新加坡) College of Computing and Data Science, Nanyang Technological University, Singapore(计算与数据科学学院,南洋理工大学,新加坡)

专题命中 知识编辑与模型理解 :language model(title,abstract)

AI总结 HMR3D通过层次化多模态表示,结合多视图图像和文本描述,提升3D场景理解的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22664 2025-12-01 cs.CV 78%

VaMP: Variational Multi-Modal Prompt Learning for Vision-Language Models

VaMP:用于视觉-语言模型的变分多模态提示学习

Silin Cheng, Kai Han

机构 * Visual AI Lab, The University of Hong Kong(香港大学视觉人工智能实验室)

专题命中 知识编辑与模型理解 :language model(title,abstract)

AI总结 VaMP提出了一种变分多模态提示学习框架,通过实例条件提示和不确定性建模提升视觉-语言模型在少样本和领域泛化任务中的性能。

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21614 2025-11-27 q-bio.QM 78%

Automated Protein Motif Localization using Concept Activation Vectors in Protein Language Model Embedding Space

利用蛋白质语言模型嵌入空间中的概念激活向量实现蛋白质motif自动定位

Ahmad Shamail, Claire D. McWhite

专题命中 知识编辑与模型理解 :language model(title,abstract)

AI总结 本文提出利用蛋白质语言模型嵌入空间中的概念激活向量实现蛋白质motif的自动化定位,通过训练线性分类器和计算内积实现高效准确的motif识别。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18416 2025-11-25 cs.CV 78%

4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation

4D-VGGT:一种具有时空意识的通用基础模型,用于动态场景几何估计

Haonan Wang, Hanyu Zhou, Haoyue Liu, Luxin Yan

机构 * National Key Lab of Multispectral Information Intelligent Processing Technology, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(多谱信息智能处理国家实验室,人工智能与自动化学院,华中科技大学) School of Computing, National University of Singapore(计算学院,新加坡国立大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

AI总结 4D-VGGT通过分而治之的时空表示方法,提升动态场景几何估计的准确性和通用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05923 2025-11-20 cs.CV 78%

Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation

Qiming Li, Zekai Ye, Xiaocheng Feng, Weihong Zhong, Weitao Ma, Xiachong Feng

专题命中 知识编辑与模型理解 :language model(title,abstract)

Comments AAAI2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09592 2025-11-14 eess.IV q-bio.QM 78%

Segment Any Tumour: An Uncertainty-Aware Vision Foundation Model for Whole-Body Analysis

Himashi Peiris, Sizhe Wang, Gary Egan, Mehrtash Harandi, Meng Law, Zhaolin Chen

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08978 2025-11-13 cs.MM cs.CV 78%

Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding

Jingtian Ma, Jingyuan Wang, Wayne Xin Zhao, Guoping Liu, Xiang Wen

机构 * School of Computer Science and Engineering, and the MOE Engineering Research Center of Advanced Computer Application Technology, Beihang University(计算机科学与工程学院,以及教育部先进计算机应用技术工程研究中心,北京航空航天大学) School of Computer Science and Engineering, the School of Economics and Management, and the MIIT Key Laboratory of Data Intelligence and Management, Beihang University(计算机科学与工程学院,经济管理学院,以及工信部数据智能与管理重点实验室,北京航空航天大学) Gaoling School of Artificial Intelligence, Renmin University of China(中关村人工智能学院,中国人民大学) DiDi Global Inc.(滴滴出行公司)

专题命中 知识编辑与模型理解 :language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25175 2025-10-30 cs.CV 78%

Test-Time Adaptive Object Detection with Foundation Model

Yingjie Gao, Yanan Zhang, Zhi Cai, Di Huang

机构 * State Key Laboratory of Complex and Critical Software Environment, Beihang University(复杂与关键软件环境国家重点实验室,北京航空航天大学) School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院)

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19911 2025-10-28 cs.CV 78%

Attention! Your Vision Language Model Could Be Maliciously Manipulated

Xiaosen Wang, Shaokang Wang, Zhijin Ge, Yuyang Luo, Shudong Zhang

专题命中 知识编辑与模型理解 :language model(title,abstract)

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21965 2025-10-28 cs.MA 78%

LLM-augmented empirical game theoretic simulation for social-ecological systems

Jennifer Shi, Christopher K. Frantz, Christian Kimmich, Saba Siddiki, Atrisha Sarkar

专题命中 知识编辑与模型理解 :LLM(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21664 2025-10-27 cs.CV q-bio.QM 78%

Foundation Models in Dermatopathology: Skin Tissue Classification

Riya Gupta, Yiwei Zong, Dennis H. Murphree

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17603 2025-10-21 cs.CV 78%

ShapeCraft: LLM Agents for Structured, Textured and Interactive 3D Modeling

Shuyuan Zhang, Chenhan Jiang, Zuoou Li, Jiankang Deng

机构 * Imperial College London(伦敦帝国学院) Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 知识编辑与模型理解 :LLM(title,abstract)

Comments NeurIPS 2025 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13169 2025-10-21 cs.CV 78%

Generate, but Verify: Reducing Hallucination in Vision-Language Models with Retrospective Resampling

Tsung-Han Wu, Heekyung Lee, Jiaxin Ge, Joseph E. Gonzalez, Trevor Darrell, David M. Chan

机构 * UC Berkeley(加州大学伯克利分校) POSTECH

专题命中 知识编辑与模型理解 :language model(title,abstract)

Comments Accepted to NeurIPS 2025; Project Page: https://reverse-vlm.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12515 2025-10-15 eess.SP 78%

HEAR: An EEG Foundation Model with Heterogeneous Electrode Adaptive Representation

Zhige Chen, Chengxuan Qin, Wenlong You, Rui Liu, Congying Chu, Rui Yang, Kay Chen Tan, Jibin Wu

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03767 2025-10-07 cs.CV 78%

CoPA: Hierarchical Concept Prompting and Aggregating Network for Explainable Diagnosis

Yiheng Dong, Yi Lin, Xin Yang

机构 * School of Electronic Information and Communications, Huazhong University of Science and Technology, Wuhan, China(电子信息学院,华中科技大学,武汉) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology, Hong Kong, China(计算机科学与工程系,香港科技大学,香港)

专题命中 知识编辑与模型理解 :prompting(title,abstract)

Comments Accepted by MICCAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25974 2025-10-01 cs.NI cs.MA 78%

OpenID Connect for Agents (OIDC-A) 1.0: A Standard Extension for LLM-Based Agent Identity and Authorization

Subramanya Nagabhushanaradhya

专题命中 知识编辑与模型理解 :LLM(title,abstract)

Comments 10 pages, 5 tables, 2 code listings. Specification proposal available at https://github.com/subramanya1997/oidc-a/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18485 2025-09-25 q-bio.NC cs.CV 78%

Deciphering Functions of Neurons in Vision-Language Models

Jiaqi Xu, Cuiling Lan, Yan Lu

机构 * University of Science and Technology of China(中国科学技术大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 知识编辑与模型理解 :language model(title,abstract)

Comments Accepted by the 31st ACM International Conference on Multimedia (ACM MM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14257 2025-09-24 cs.CV 78%

Mitigating Hallucination in Large Vision-Language Models through Aligning Attention Distribution to Information Flow

Jianfei Zhao, Feng Zhang, Xin Sun, Chong Feng

机构 * School of Computer Science and Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学) Zhongguancun Academy(中关村学院) Southeast Academy of Information Technology, Beijing Institute of Technology(信息技术东南学院,北京理工大学)

专题命中 知识编辑与模型理解 :language model(title,abstract)

Comments Accepted to Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15416 2025-09-22 cs.CV 78%

NeuroRAD-FM: A Foundation Model for Neuro-Oncology with Distributionally Robust Training

Moinak Bhattacharya, Angelica P. Kurtz, Fabio M. Iwamoto, Prateek Prasanna, Gagandeep Singh

机构 * Department of Biomedical Informatics, Stony Brook University(生物医学信息学系,石溪大学) Department of Radiology, Columbia University Irving Medical Center(放射学系,哥伦比亚大学伊万杰琳医疗中心) Department of Neuro-Oncology, Columbia University Irving Medical Center(神经肿瘤学系,哥伦比亚大学伊万杰琳医疗中心)

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00314 2025-09-03 eess.SP 78%

CoMET: A Contrastive-Masked Brain Foundation Model for Universal EEG Representation

Ang Li, Zikai Wang, Liuyin Yang, Zhenyu Wang, Tianheng Xu, Honglin Hu, Marc M. Van Hulle

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16207 2025-08-26 cs.CV 78%

T-MASK: Temporal Masking for Probing Foundation Models across Camera Views in Driver Monitoring

Thinesh Thiyakesan Ponbagavathi, Kunyu Peng, Alina Roitberg

机构 * Institute of Artificial Intelligence(人工智能研究所) University of Stuttgart(斯图加特大学) Institute for Anthropomatics(人机学研究所) Karlsruhe Institute of Technology(卡尔斯鲁厄技术大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

Comments This paper has been accepted by 26th IEEE International Conference on Intelligent Transportation Systems ITSC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏