arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-23 至 2025-10-23 共收录 5 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 5 篇

2509.18582 2025-10-23 cs.CV 79%

The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers

Daiqing Qi, Handong Zhao, Jing Shi, Simon Jenni, Yifei Fan, Franck Dernoncourt, Scott Cohen, Sheng Li

机构 * University of Virginia(弗吉尼亚大学) Adobe(Adobe公司)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);分类 cs.CV

Journal ref CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12718 2025-10-23 cs.CV cs.MM 79%

ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding

Zhenxing Zhang, Yaxiong Wang, Lechao Cheng, Zhun Zhong, Dan Guo, Meng Wang

机构 * School of Computer Science and Information Engineering, Hefei University of Technology, China(计算机科学与信息工程学院,合肥工业大学,中国) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, China(人工智能研究所,合肥综合性国家科学中心,中国)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19599 2025-10-23 cs.CV cs.AI 79%

XBench: A Comprehensive Benchmark for Visual-Language Explanations in Chest Radiography

Haozhe Luo, Shelley Zixin Shu, Ziyu Zhou, Sebastian Otalora, Mauricio Reyes

机构 * ARTORG Center for Biomedical Engineering Research, University of Bern, Switzerland(ARTORG生物医学工程研究中心,伯尔尼大学,瑞士) Shanghai Jiao Tong University, China(上海交通大学,中国) Kaiko.AI, Switzerland(Kaiko.AI,瑞士) Dept. of Radiation Oncology, Inselspital, Bern University Hospital(放射肿瘤科,因斯普尔茨医院,伯尔尼大学医院)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19574 2025-10-23 cs.CV cs.CR 57%

Can You Trust What You See? Alpha Channel No-Box Attacks on Video Object Detection

Ariana Yi, Ce Zhou, Liyang Xiao, Qiben Yan

机构 * Mission San Jose High School(Mission San Jose 高中) Missouri University of Science and Technology(密苏里科学与技术大学) Michigan State University(密歇根州立大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19878 2025-10-23 cs.CL cs.IR 50%

CausalRAG: Integrating Causal Graphs into Retrieval-Augmented Generation

Nengbo Wang, Xiaotian Han, Jagdip Singh, Jing Ma, Vipin Chaudhary

机构 * Department of Computer and Data Sciences, Case Western Reserve University(计算机与数据科学系,凯斯西储大学) Department of Design and Innovation, Case Western Reserve University(设计与创新系,凯斯西储大学)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted at Findings of ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏