arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 7583 信号源:cs.CL, cs.AI, cs.LG

1. 知识编辑与模型理解 7583 篇

2506.04562 2025-08-22 cs.GR cs.CV 78%

Handle-based Mesh Deformation Guided By Vision Language Model

Xingpeng Sun, Shiyang Jia, Zherong Pan, Kui Wu, Aniket Bera

机构 * Purdue University(普渡大学) LightSpeed Studios University of California San Diego(加州大学圣地亚哥分校)

专题命中 知识编辑与模型理解 :language model(title,abstract)

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12586 2025-08-19 cs.CV 78%

Foundation Model for Skeleton-Based Human Action Understanding

Hongsong Wang, Wanjiang Weng, Junbo Wang, Fang Zhao, Guo-Sen Xie, Xin Geng, Liang Wang

机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及交叉应用关键实验室(东南大学),中华人民共和国教育部,中国) School of Software, Northwestern Polytechnical University(西北工业大学软件学院) State Key Laboratory for Novel Software Technology and School of Intelligence Science and Technology, Nanjing University(新型软件技术国家重点实验室和南京大学智能科学与技术学院) School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院)

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

Comments Accepted by TPAMI, Code is available at: https://github.com/wengwanjiang/FoundSkelModel

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10357 2025-08-19 cs.CV 78%

Optimization of Prompt Learning via Multi-Knowledge Representation for Vision-Language Models

Enming Zhang, Bingke Zhu, Yingying Chen, Qinghai Miao, Ming Tang, Jinqiao Wang

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) Wuhan AI Research(武汉人工智能研究所) Peng Cheng Laboratory(鹏城实验室)

专题命中 知识编辑与模型理解 :language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09184 2025-07-24 cs.CV 78%

MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models

Qiyan Zhao, Xiaofeng Zhang, Yiheng Li, Yun Xing, Xiaosong Yuan, Feilong Tang, Sinan Fan, Xuhang Chen, Xuyao Zhang, Dahan Wang

机构 * FKLPRIU, Xiamen University of Technology, China(福克斯理工国际大学,厦门理工学院,中国) Shanghai Jiao Tong University, China(上海交通大学,中国) Nanyang Technological University, Singapore(南洋理工大学,新加坡) Jilin University, China(吉林大学,中国) Monash University, Australia(墨尔本大学,澳大利亚) Zhejiang University, China(浙江大学,中国) Huizhou University, China(惠州大学,中国) Chinese Academy of Sciences, China(中国科学院,中国)

专题命中 知识编辑与模型理解 :language model(title,abstract)

Comments Accepted in ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05887 2025-07-21 cs.CV 78%

GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing

Xianzhi Ma, Jianhui Li, Changhua Pei, Hao Liu

机构 * Institute of Space Earth Science, School of Frontier Sciences, Nanjing University(南京大学空间地球科学学院) Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心)

专题命中 知识编辑与模型理解 :language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12938 2025-07-18 eess.IV cs.CV 78%

Unleashing Vision Foundation Models for Coronary Artery Segmentation: Parallel ViT-CNN Encoding and Variational Fusion

Caixia Dong, Duwei Dai, Xinyi Han, Fan Liu, Xu Yang, Zongfang Li, Songhua Xu

机构 * National-Local Joint Engineering Research Center of Biodiagnosis \& Biotherapy, the Second Affiliated Hospital of Xi’an Jiaotong University, Xi’an, 710004, China Institute of Medical Artificial Intelligence, the Second Affiliated Hospital of Xi’an Jiaotong University, Xi'an, 710004, China Viadrina European University, Frankfurt (Oder), 15230, Germany

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

Journal ref MICCAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18675 2025-07-16 cs.CV 78%

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models

Pooyan Rahmanzadehgervi, Hung Huy Nguyen, Rosanne Liu, Long Mai, Anh Totti Nguyen

机构 * Auburn University(奥本大学) Google DeepMind, ML Collective(谷歌DeepMind及ML集体) Adobe Research(Adobe研究)

专题命中 知识编辑与模型理解 :language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06418 2025-07-10 q-bio.QM cs.CV stat.AP 78%

PAST: A multimodal single-cell foundation model for histopathology and spatial transcriptomics in cancer

Changchun Yang, Haoyang Li, Yushuai Wu, Yilan Zhang, Yifeng Jiao, Yu Zhang, Rihan Huang, Yuan Cheng, Yuan Qi, Xin Guo, Xin Gao

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04118 2025-07-08 cs.CV 78%

PromptSR: Cascade Prompting for Lightweight Image Super-Resolution

Wenyang Liu, Chen Cai, Jianjun Gao, Kejun Wu, Yi Wang, Kim-Hui Yap, Lap-Pui Chau

机构 * School of Electrical and Electronics Engineering, Nanyang Technological University, Singapore(南洋理工大学电子与电气工程学院) School of Electronic Information and Communications, Huazhong University of Science and Technology, Wuhan 430074, China(华中科技大学电子信息与通信学院) Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University, Hong Kong(香港理工大学电子与电气工程系)

专题命中 知识编辑与模型理解 :prompting(title,abstract)

Comments Accepted in TMM

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11620 2025-07-02 cs.SE 78%

Assessing Correctness in LLM-Based Code Generation via Uncertainty Estimation

Arindam Sharma, Cristina David

专题命中 知识编辑与模型理解 :LLM(title,abstract)

Comments 18 pages and 3 References Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23590 2025-07-01 cs.CV 78%

CAI: Caption-Sensitive Attention Intervention for Mitigating Object Hallucination in Large Vision-Language Models

Qiming Li, Zekai Ye, Xiaocheng Feng, Weihong Zhong, Libo Qin, Ruihan Chen, Baohang Li, Kui Jiang, Yaowei Wang, Ting Liu, Bing Qin

专题命中 知识编辑与模型理解 :language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22979 2025-07-01 cs.CV 78%

Probabilistic Prototype Calibration of Vision-Language Models for Generalized Few-shot Semantic Segmentation

Jie Liu, Jiayi Shen, Pan Zhou, Jan-Jakob Sonke, Efstratios Gavves

机构 * University of Amsterdam(阿姆斯特丹大学) Singapore Management University(新加坡管理大学) The Netherlands Cancer Institute(荷兰癌症研究院)

专题命中 知识编辑与模型理解 :language model(title,abstract)

Comments ICCV2025 Proceeding

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15866 2025-06-23 cs.SI 78%

Understanding Online Polarization Through Human-Agent Interaction in a Synthetic LLM-Based Social Network

Tim Donkers, Jürgen Ziegler

专题命中 知识编辑与模型理解 :LLM(title,abstract)

Comments Accepted for publication in the Proceedings of the Nineteenth International AAAI Conference on Web and Social Media (ICWSM 2025). This is the authors' version of the work, with corrections to table cross-references. The definitive Version of Record is available at https://doi.org/10.1609/icwsm.v19i1.35826. arXiv admin note: substantial text overlap with arXiv:2502.01340

Journal ref Proceedings of the Nineteenth International AAAI Conference on Web and Social Media (ICWSM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02258 2025-06-04 eess.AS cs.SD 78%

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition?

Mohd Mujtaba Akhtar, Orchid Chetia Phukan, Girish, Swarup Ranjan Behera, Ananda Chandra Nayak, Sanjib Kumar Nayak, Arun Balaji Buduru, Rajesh Sharma

机构 * V.B.S.P.U, India(印度V.B.S.P.U大学) IIIT-Delhi, India(印度德里印度理工学院) UPES, India(印度UPES大学) KAC, India(印度KAC机构) VSSUT, India(印度VSSUT大学) University of Tartu, Estonia(爱沙尼亚塔尔图大学) Plaksha University,India(印度Plaksha大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

Comments Accepted to EUSIPCO 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01867 2025-06-03 q-bio.NC eess.SP 78%

EEG Foundation Models for BCI Learn Diverse Features of Electrophysiology

Mattson Ogg, Rahul Hingorani, Diego Luna, Griffin W. Milsap, William G. Coon, Clara A. Scholl

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

Comments Two figures, one table, six pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01203 2025-06-03 cs.CV 78%

Self-Supervised Multi-View Representation Learning using Vision-Language Model for 3D/4D Facial Expression Recognition

Muzammil Behzad

机构 * King Fahd University of Petroleum and Minerals(国王法赫德石油矿物大学)

专题命中 知识编辑与模型理解 :language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23358 2025-05-30 cs.CV 78%

Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model

Reem AlJunaid, Muzammil Behzad

机构 * KFUPM(科威特理工学院)

专题命中 知识编辑与模型理解 :language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13746 2025-05-21 cs.CV eess.IV 78%

ReSW-VL: Representation Learning for Surgical Workflow Analysis Using Vision-Language Model

Satoshi Kondo

机构 * Muroran Institute of Technology(茂兰技术学院)

专题命中 知识编辑与模型理解 :language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12233 2025-05-20 eess.IV cs.CV 78%

PRETI: Patient-Aware Retinal Foundation Model via Metadata-Guided Representation Learning

Yeonkyung Lee, Woojung Han, Youngjun Jun, Hyeonmin Kim, Jungkyung Cho, Seong Jae Hwang

机构 * Yonsei University(延世大学) Mediwhale(Mediwhale公司)

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

Comments MICCAI2025 early accept

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09336 2025-05-15 cs.CV 78%

Unsupervised Multiview Contrastive Language-Image Joint Learning with Pseudo-Labeled Prompts Via Vision-Language Model for 3D/4D Facial Expression Recognition

Muzammil Behzad

专题命中 知识编辑与模型理解 :language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11999 2025-04-17 cs.CV 78%

A Complex-valued SAR Foundation Model Based on Physically Inspired Representation Learning

Mengyu Wang, Hanbo Bi, Yingchao Feng, Linlin Xin, Shuo Gong, Tianqi Wang, Zhiyuan Yan, Peijin Wang, Wenhui Diao, Xian Sun

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05445 2025-04-09 cs.HC 78%

Probing the Visualization Literacy of Vision Language Models: the Good, the Bad, and the Ugly

Lianghan Dong, Anamaria Crisan

专题命中 知识编辑与模型理解 :language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01009 2025-04-02 cs.CV 78%

GECKO: Gigapixel Vision-Concept Contrastive Pretraining in Histopathology

Saarthak Kapse, Pushpak Pati, Srikar Yellapragada, Srijan Das, Rajarsi R. Gupta, Joel Saltz, Dimitris Samaras, Prateek Prasanna

专题命中 知识编辑与模型理解 :pretraining(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15356 2025-04-02 cs.CV 78%

Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced Optimization

Xinyu Lyu, Beitao Chen, Lianli Gao, Jingkuan Song, Heng Tao Shen

专题命中 知识编辑与模型理解 :language model(title,abstract)

Comments Accepted by NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18290 2025-03-27 cs.CV 78%

Stealthy Backdoor Attack in Self-Supervised Learning Vision Encoders for Large Vision Language Models

Zhaoyi Liu, Huan Zhang

专题命中 知识编辑与模型理解 :language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15839 2025-03-18 cs.CV 78%

VaLiD: Mitigating the Hallucination of Large Vision Language Models by Visual Layer Fusion Contrastive Decoding

Jiaqi Wang, Yifei Gao, Jitao Sang

专题命中 知识编辑与模型理解 :language model(title,abstract)

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04992 2025-02-27 eess.IV cs.CV 78%

VisionFM: a Multi-Modal Multi-Task Vision Foundation Model for Generalist Ophthalmic Artificial Intelligence

Jianing Qiu, Jian Wu, Hao Wei, Peilun Shi, Minqing Zhang, Yunyun Sun, Lin Li, Hanruo Liu, Hongyi Liu, Simeng Hou, Yuyang Zhao, Xuehui Shi, Junfang Xian, Xiaoxia Qu, Sirui Zhu, Lijie Pan, Xiaoniao Chen, Xiaojia Zhang, Shuai Jiang, Kebing Wang, Chenlong Yang, Mingqiang Chen, Sujie Fan, Jianhua Hu, Aiguo Lv, Hui Miao, Li Guo, Shujun Zhang, Cheng Pei, Xiaojuan Fan, Jianqin Lei, Ting Wei, Junguo Duan, Chun Liu, Xiaobo Xia, Siqi Xiong, Junhong Li, Benny Lo, Yih Chung Tham, Tien Yin Wong, Ningli Wang, Wu Yuan

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

Journal ref The latest VisionFM work has been published in NEJM AI, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00095 2025-01-20 cs.CV 78%

OPCap:Object-aware Prompting Captioning

Feiyang Huang

专题命中 知识编辑与模型理解 :prompting(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.17345 2025-01-16 cs.CV 78%

3VL: Using Trees to Improve Vision-Language Models' Interpretability

Nir Yellinek, Leonid Karlinsky, Raja Giryes

专题命中 知识编辑与模型理解 :language model(title,abstract)

Comments accepted to IEEE TIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04815 2025-01-10 cs.CV 78%

Towards Generalizable Trajectory Prediction Using Dual-Level Representation Learning And Adaptive Prompting

Kaouther Messaoud, Matthieu Cord, Alexandre Alahi

专题命中 知识编辑与模型理解 :prompting(title);pretraining(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏