arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-12 至 2025-09-12 共收录 4 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 4 篇

2509.09584 2025-09-12 cs.CV cs.RO 79%

Visual Grounding from Event Cameras

Lingdong Kong, Dongyue Lu, Ao Liang, Rong Li, Yuhao Dong, Tianshuai Hu, Lai Xing Ng, Wei Tsang Ooi, Benoit R. Cottereau

机构 * NUS(新加坡国立大学) HKUST(GZ)(香港科技大学(广州)) NTU(南洋理工大学) HKUST(香港科技大学) I 2 R, A*STAR(新加坡科技研究局) IPAL, CNRS(法国国家科学研究中心IPAL) CerCo, CNRS(法国国家科学研究中心CerCo)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Abstract Paper (Non-Archival) @ ICCV 2025 NeVi Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09014 2025-09-12 cs.CV cs.CL 70%

COCO-Urdu: A Large-Scale Urdu Image-Caption Dataset with Multimodal Quality Estimation

Umair Hassan

机构 * Independent Researcher(独立研究者)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments 17 pages, 3 figures, 3 tables. Dataset available at https://huggingface.co/datasets/umairhassan02/urdu-translated-coco-captions-subset. Scripts and notebooks to reproduce results available at https://github.com/umair-hassan2/COCO-Urdu

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09172 2025-09-12 cs.CV 57%

Bridging the Gap Between Ideal and Real-world Evaluation: Benchmarking AI-Generated Image Detection in Challenging Scenarios

Chunxiao Li, Xiaoxiao Wang, Meiling Li, Boming Miao, Peng Sun, Yunjian Zhang, Xiangyang Ji, Yao Zhu

机构 * Beijing Normal University(北京师范大学) University of Chinese Academy of Sciences(中国科学院大学) Fudan University(复旦大学) Central University of Finance and Economics(中央财经大学) Tsinghua University(清华大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09281 2025-09-12 cs.HC 50%

Flip Co-op: Cooperative Takeovers in Shared Autonomy

Sandeep Banik, Naira Hovakimyan

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 11 pages and 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏