arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-05 至 2025-09-05 共收录 10 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 10 篇

2509.04243 2025-09-05 cs.CV cs.AI 84%

Learning Active Perception via Self-Evolving Preference Optimization for GUI Grounding

Wanfu Wang, Qipeng Huang, Guangquan Xue, Xiaobo Liang, Juntao Li

机构 * Wanfu Wang, Qipeng Huang, Guangquan Xue, Xiaobo Liang, Juntao Li(作者)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03800 2025-09-05 cs.CV 83%

MedVista3D: Vision-Language Modeling for Reducing Diagnostic Errors in 3D CT Disease Detection, Understanding and Reporting

Yuheng Li, Yenho Chen, Yuxiang Lai, Jike Zhong, Vanessa Wildman, Xiaofeng Yang

机构 * Department of Biomedical Engineering(生物医学工程系) Georgia Institute of Technology(佐治亚理工学院) Department of Machine Learning(机器学习系) Department of Radiation Oncology(放射肿瘤科) Emory University School of Medicine(埃默里大学医学院) University of Southern California(南加州大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14904 2025-09-05 cs.CV cs.AI 81%

TriCLIP-3D: A Unified Parameter-Efficient Framework for Tri-Modal 3D Visual Grounding based on CLIP

Fan Li, Zanyi Wang, Zeyi Huang, Guang Dai, Jingdong Wang, Mengmeng Wang

机构 * Xi’an Jiaotong University(西安交通大学) SGIT AI Lab(SGIT人工智能实验室) Zhejiang University of Technology(浙江工业大学) Huawei(华为)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03961 2025-09-05 cs.CV cs.AI 73%

Multimodal Feature Fusion Network with Text Difference Enhancement for Remote Sensing Change Detection

Yijun Zhou, Yikui Zhai, Zilu Ying, Tingfeng Xian, Wenlve Zhou, Zhiheng Zhou, Xiaolin Tian, Xudong Jia, Hongsheng Zhang, C. L. Philip Chen

机构 * College of Electronics and Information Engineering, Wuyi University(威怡大学电子与信息工程学院) School of Electronic and Information Engineering and the Key Laboratory of Big Data and Intelligent Robot, Ministry of Education, South China University of Technology(电子与信息工程学院和大数据与智能机器人重点实验室,华南理工大学) State Key Laboratory of Lunar and Planetary Sciences, Macau University of Science and Technology(澳门大学地球和行星科学国家重点实验室) College of Engineering and Computer Science, California State University, Northridge(工程与计算机科学学院,加州大学北岭分校) Department of Geography, The University of Hong Kong(地理系,香港大学) Faculty of Computer Science and Engineering, S(计算机科学与工程学院,S)

专题命中 视觉定位与Grounding :vision language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04326 2025-09-05 cs.CV 70%

Efficient Odd-One-Out Anomaly Detection

Silvio Chito, Paolo Rabino, Tatiana Tommasi

机构 * Politecnico di Torino(托斯尼亚理工学院)

专题命中 视觉定位与Grounding :visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV

Comments Accepted at ICIAP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04180 2025-09-05 cs.CV cs.AI 62%

VisioFirm: Cross-Platform AI-assisted Annotation Tool for Computer Vision

Safouane El Ghazouali, Umberto Michelucci

机构 * TOELT LLC AI lab(TOELT LLC人工智能实验室)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03793 2025-09-05 cs.MA cs.AI 57%

SAMVAD: A Multi-Agent System for Simulating Judicial Deliberation Dynamics in India

Prathamesh Devadiga, Omkaar Jayadev Shetty, Pooja Agarwal

机构 * PES University(PES大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03626 2025-09-05 cs.AI 57%

Explainable Knowledge Graph Retrieval-Augmented Generation (KG-RAG) with KG-SMILE

Zahra Zehtabi Sabeti Moghaddam, Zeinab Dehghani, Maneeha Rani, Koorosh Aslansefat, Bhupesh Kumar Mishra, Rameez Raja Kureshi, Dhavalkumar Thakker

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04254 2025-09-05 cs.HC 50%

MuMTAffect: A Multimodal Multitask Affective Framework for Personality and Emotion Recognition from Physiological Signals

Meisam Jamshidi Seikavandi, Fabricio Batista Narcizo, Ted Vucurevich, Andrew Burke Dittberner, Paolo Burelli

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12065 2025-09-05 cs.CL cs.FL 50%

Autoformalization in the Wild: Assessing LLMs on Real-World Mathematical Definitions

Lan Zhang, Marco Valentino, Andre Freitas

机构 * Department of Computer Science, University of Manchester(曼彻斯特大学计算机科学系) School of Computer Science, University of Sheffield(谢菲尔德大学计算机科学学院) Idiap Research Institute(Idiap研究机构) National Biomarker Centre, CRUK Manchester Institute(国家生物标志物中心、CRUK曼彻斯特研究所)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments EMNLP 2025 Camera-Ready Version

详情

展开后加载摘要…

URL PDF HTML 收藏