arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-11-04 至 2025-11-04 共收录 76 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 16 篇

2511.00269 2025-11-04 cs.CV cs.AI 62%

FedReplay: A Feature Replay Assisted Federated Transfer Learning Framework for Efficient and Privacy-Preserving Smart Agriculture

Long Li, Jiajia Li, Dong Chen, Lina Pu, Haibo Yao, Yanbo Huang

机构 * Department of Electrical and Computer Engineering, The University of Alabama(电气与计算机工程系,阿拉巴马大学) Electrical and Computer Engineering, Michigan State University(电气与计算机工程,密歇根州立大学) Agricultural and Biological Engineering, Mississippi State University(农业与生物工程,密苏里州立大学) Department of Computer Science, University of Alabama(计算机科学系,阿拉巴马大学) USDA-ARS Genetics and Sustainbale Agriculture(美国农业部ARS基因与可持续农业)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10821 2025-11-04 cs.CV cs.AI cs.CL 62%

VideoExplorer: Think With Videos For Agentic Long-Video Understanding

Huaying Yuan, Zheng Liu, Junjie Zhou, Hongjin Qian, Yan Shu, Nicu Sebe, Ji-Rong Wen, Zhicheng Dou

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Beijing University of Posts and Telecommunications(北京邮电大学) Hong Kong Polytechnic University(香港理工大学) Peking University(北京大学) University of Trento(特伦托大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01753 2025-11-04 cs.LO cs.AI cs.PL 57%

SM-based Semantics for Answer Set Programs Containing Conditional Literals and Arithmetic

Zachary Hansen, Yuliya Lierler

机构 * University of Nebraska Omaha(内布拉斯加大学奥马哈分校)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments This version corrects the review of tau for negated atoms, and clarifies the distinction between global and local variables in conditional literals (the supporting proofs are also updated accordingly)

Journal ref In Practical Aspects of Declarative Languages: 27th International Symposium, PADL 2025, Denver, CO, USA, January 20-21, 2025, Proceedings. Springer-Verlag, Berlin, Heidelberg, 71-87

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16800 2025-11-04 cs.CV 57%

Phys4DGen: Physics-Compliant 4D Generation with Multi-Material Composition Perception

Jiajing Lin, Zhenzhong Wang, Dejun Xu, Shu Jiang, YunPeng Gong, Min Jiang

机构 * School of Informatics, Xiamen University(厦门大学信息学院)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

Comments Accepted by ACM MM 2025. Project Page: https://jiajinglin.github.io/Phys4DGen

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16093 2025-11-04 cs.CL cs.AI 57%

Beyond Pointwise Scores: Decomposed Criteria-Based Evaluation of LLM Responses

Fangyi Yu, Nabeel Seedat, Dasha Herrmannova, Frank Schilder, Jonathan Richard Schwarz

机构 * Thomson Reuters Foundational Research(汤姆森·路透基础研究) Thomson Reuters Labs(汤姆森·路透实验室) Imperial College London(帝国理工学院伦敦分校)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted by 2025 EMNLP industry track

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01390 2025-11-04 cs.CY cs.AI 57%

Recognising, Anticipating, and Mitigating LLM Pollution of Online Behavioural Research

Raluca Rilla, Tobias Werner, Hiromu Yakura, Iyad Rahwan, Anne-Marie Nussberger

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08677 2025-11-04 quant-ph physics.atom-ph 50%

Multiparameter estimation with an array of entangled atomic sensors

Yifan Li, Lex Joosten, Youcef Baamara, Paolo Colciaghi, Alice Sinatra, Philipp Treutlein, Tilman Zibold

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05049 2025-11-04 cs.SI cs.CY 50%

Uncovering the Sociodemographic Fabric of Reddit

Federico Cinus, Corrado Monti, Paolo Bajardi, Gianmarco De Francisci Morales

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 文档图表理解 1 篇

2502.01341 2025-11-04 cs.CL 50%

AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding

Ahmed Masry, Juan A. Rodriguez, Tianyu Zhang, Suyuchen Wang, Chao Wang, Aarash Feizi, Akshay Kalkunte Suresh, Abhay Puri, Xiangru Jian, Pierre-André Noël, Sathwik Tejaswi Madhusudhan, Marco Pedersoli, Bang Liu, Nicolas Chapados, Yoshua Bengio, Enamul Hoque, Christopher Pal, Issam H. Laradji, David Vazquez, Perouz Taslakian, Spandana Gella, Sai Rajeswar

机构 * ServiceNow York University(约克大学) Mila – Quebec AI Institute(魁北克人工智能研究院) École de Technologie Supérieure(魁北克高等技术学院) Université de Montréal(蒙特利尔大学) McGill University(麦吉尔大学) University of Waterloo(滑铁卢大学) CIFAR AI Chair(CIFAR人工智能 chair) Polytechnique Montréal(蒙特利尔理工学院) University of British Columbia(不列颠哥伦比亚大学)

专题命中 文档图表理解 :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. GUI与屏幕智能体 3 篇

2510.27255 2025-11-04 cs.CV 57%

Enhancing Spatio-Temporal Zero-shot Action Recognition with Language-driven Description Attributes

Yehna Kim, Young-Eun Kim, Seong-Whan Lee

机构 * Department of Artificial Intelligence, Korea University(韩国大学人工智能系)

专题命中 GUI与屏幕智能体 :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00933 2025-11-04 cs.RO cs.CV 57%

Fast-SmartWay: Panoramic-Free End-to-End Zero-Shot Vision-and-Language Navigation

Xiangyu Shi, Zerui Li, Yanyuan Qiao, Qi Wu

机构 * Australian Institute for Machine Learning, the University of Adelaide(澳大利亚机器学习研究所、阿德莱德大学) CREATE Lab, Swiss Federal Institute of Technology Lausanne (EPFL)(洛桑联邦理工学院CREATE实验室)

专题命中 GUI与屏幕智能体 :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23763 2025-11-04 cs.RO cs.CL cs.CV 57%

RoboOmni: Proactive Robot Manipulation in Omni-modal Context

Siyin Wang, Jinlan Fu, Feihong Liu, Xinzhe He, Huangxuan Wu, Junhao Shi, Kexin Huang, Zhaoye Fei, Jingjing Gong, Zuxuan Wu, Yu-Gang Jiang, See-Kiong Ng, Tat-Seng Chua, Xipeng Qiu

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) National University of Singapore(新加坡国立大学)

专题命中 GUI与屏幕智能体 :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与鲁棒性 11 篇

2504.08809 2025-11-04 cs.LG 83%

Decoupling Contrastive Decoding: Robust Hallucination Mitigation in Multimodal Large Language Models

Wei Chen, Xin Yan, Bin Wen, Fan Yang, Tingting Gao, Di Zhang, Long Chen

机构 * HKUST(香港科技大学) University of Waterloo(滑铁卢大学) Kuaishou Technology(快手科技)

专题命中 幻觉与鲁棒性 :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.LG

Comments 17 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07409 2025-11-04 cs.CV cs.LG 81%

MGPATH: Vision-Language Model with Multi-Granular Prompt Learning for Few-Shot WSI Classification

Anh-Tien Nguyen, Duy Minh Ho Nguyen, Nghiem Tuong Diep, Trung Quoc Nguyen, Nhat Ho, Jacqueline Michelle Metsch, Miriam Cindy Maurer, Daniel Sonntag, Hanibal Bohnenberger, Anne-Christin Hauschild

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV、cs.LG

Comments Published in Transactions on Machine Learning Research (09/2025)

Journal ref Transactions on Machine Learning Research (09/2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00509 2025-11-04 cs.AI cs.CR 70%

Reimagining Safety Alignment with An Image

Yifan Xia, Guorui Chen, Wenqian Yu, Zhijiang Li, Philip Torr, Jindong Gu

机构 * School of Information Management, Wuhan University, Wuhan, China(武汉大学信息管理学院) Torr Vision Group, University of Oxford, Oxford, United Kingdom(牛津大学)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00411 2025-11-04 cs.LG cs.AI cs.CV 67%

Enhancing Adversarial Transferability by Balancing Exploration and Exploitation with Gradient-Guided Sampling

Zenghao Niu, Weicheng Xie, Siyang Song, Zitong Yu, Feng Liu, Linlin Shen

机构 * School of Computer Science & Software Engineering, Shenzhen University, China(深圳大学计算机科学与软件工程学院) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ), Shenzhen, China(广东省人工智能与数字经济发展实验室(深圳)) Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen University, China(广东省智能信息处理省级重点实验室) School of Computer Science, University of Exeter, U.K.(埃克塞特大学计算机科学学院) Department of Computing and Information Technology, Great Bay University, China(大鹏大学计算与信息技术系) Computer Vision Institute, School of Artificial Intelligence, Shenzhen University, China(人工智能学院计算机视觉研究所)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments accepted by iccv 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26466 2025-11-04 cs.CV cs.LG 62%

Representation-Level Counterfactual Calibration for Debiased Zero-Shot Recognition

Pei Peng, MingKun Xie, Hang Hao, Tong Jin, ShengJun Huang

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00831 2025-11-04 cs.CV cs.AI 62%

Enhancing Adversarial Transferability in Visual-Language Pre-training Models via Local Shuffle and Sample-based Attack

Xin Liu, Aoyang Zhou, Aoyang Zhou

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Accepted by NAACL2025 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00446 2025-11-04 cs.CV cs.CR cs.LG 62%

ToxicTextCLIP: Text-Based Poisoning and Backdoor Attacks on CLIP Pre-training

Xin Yao, Haiyang Zhao, Yimin Chen, Jiawei Guo, Kecheng Huang, Ming Zhao

机构 * School of Computer Science and Engineering, Central South University(中南大学计算机科学与工程学院) Miner School of Computer & Information Sciences, University of Massachusetts Lowell(马萨诸塞大学洛厄尔分校Miner计算机与信息科学学院)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.LG

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01438 2025-11-04 cs.LG 57%

The Curvature Rate λ: A Scalar Measure of Input-Space Sharpness in Neural Networks

Jacob Poschl

机构 * University of California, Santa Cruz(加州大学圣克ruz分校)

专题命中 幻觉与鲁棒性 :grounding(abstract);分类 cs.LG

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00523 2025-11-04 cs.CV 57%

SegDebias: Test-Time Bias Mitigation for ViT-Based CLIP via Segmentation

Fangyu Wu, Yujun Cai

专题命中 幻觉与鲁棒性 :vision language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23945 2025-11-04 cs.CL cs.AI 57%

A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models

Sriram Balasubramanian, Samyadeep Basu, Soheil Feizi

机构 * Department of Computer Science University of Maryland, College Park(计算机科学系马里兰大学 College Park)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.AI

Comments Accepted in EMNLP 2025, 34 pages, 25 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00776 2025-11-04 cs.SE 50%

A Systematic Literature Review of Code Hallucinations in LLMs: Characterization, Mitigation Methods, Challenges, and Future Directions for Reliable AI

Cuiyun Gao, Guodong Fan, Chun Yong Chong, Shizhan Chen, Chao Liu, David Lo, Zibin Zheng, Qing Liao

专题命中 幻觉与鲁棒性 :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. VLM训练与架构 15 篇

2504.00502 2025-11-04 cs.CV cs.CL 85%

ShortV: Efficient Multimodal Large Language Models by Freezing Visual Tokens in Ineffective Layers

Qianhao Yuan, Qingyu Zhang, Yanjiang Liu, Jiawei Chen, Yaojie Lu, Hongyu Lin, Jia Zheng, Xianpei Han, Le Sun

机构 * Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 VLM训练与架构 :multimodal large language model(title,abstract);LLaVA(abstract);MLLM(abstract);分类 cs.CV

Comments Published as a conference paper at ICCV 2025. Project page: https://github.com/icip-cas/ShortV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00821 2025-11-04 cs.CV 85%

OMEGA: Optimized Multimodal Position Encoding Index Derivation with Global Adaptive Scaling for Vision-Language Models

Ruoxiang Huang, Xindian Ma, Rundong Kong, Zhen Yuan, Peng Zhang

机构 * Tianjin University(天津大学)

专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);LLaVA(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07416 2025-11-04 cs.LG cs.AI 84%

LiteVLM: A Low-Latency Vision-Language Model Inference Pipeline for Resource-Constrained Environments

Jin Huang, Yuchao Jin, Le An, Josh Park

机构 * NVIDIA(英伟达)

专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00120 2025-11-04 cs.CV cs.AI 76%

VLM6D: VLM based 6Dof Pose Estimation based on RGB-D Images

Md Selim Sarowar, Sungho Kim

专题命中 VLM训练与架构 :VLM(title);分类 cs.CV、cs.AI

Comments This paper has been accepted to IEIE( The Institute Of Electronics and Information Engineering, South Korea) Fall,2025 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21344 2025-11-04 cs.CV cs.AI q-bio.QM 76%

Vision-Language Model-Based Semantic-Guided Imaging Biomarker for Lung Nodule Malignancy Prediction

Luoting Zhuang, Seyed Mohammad Hossein Tabatabaei, Ramin Salehi-Rad, Linh M. Tran, Denise R. Aberle, Ashley E. Prosper, William Hsu

机构 * organization= Medical \& Imaging Informatics, Department of Radiological Sciences, David Geffen School of Medicine at UCLA , city= Los Angeles , postcode= 90095 , state= CA , country= USA organization= Department of Medicine, Division of Pulmonology Critical Care, David Geffen School of Medicine at UCLA , city= Los Angeles , postcode= 90095 , state= CA , country= USA

专题命中 VLM训练与架构 :vision-language model(title);分类 cs.CV、cs.AI

Journal ref Journal of Biomedical Informatics 172 (2025) 104947

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01082 2025-11-04 cs.CV cs.AI cs.LG 75%

GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction

Narges Ghasemi, Amir Ziashahabi, Salman Avestimehr, Cyrus Shahabi

机构 * of Computer Science, University of Southern California, Los Angeles, CA, USA Computer Engineering, University of Southern California, Los Angeles, CA, USA

专题命中 VLM训练与架构 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted to IEEE International Conference on Data Mining (ICDM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03724 2025-11-04 cs.CV cs.CL 74%

CrowdVLM-R1: Expanding R1 Ability to Vision Language Model for Crowd Counting using Fuzzy Group Relative Policy Reward

Zhiqiang Wang, Pengbin Feng, Yanbin Lin, Shuzhang Cai, Zongao Bian, Jinghua Yan, Xingquan Zhu

机构 * Florida Atlantic University(佛罗里达大学) University of Southern California(南加州大学) University of Texas at Dallas(德克萨斯大学达拉斯分校) Georgia Institute of Technology(佐治亚理工学院) University of Utah(犹他大学)

专题命中 VLM训练与架构 :vision language model(title);分类 cs.CV

Comments 10 pages, 6 figures and 4 tables

Journal ref 2025 IEEE International Conference on Big Data (IEEE BigData 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏