arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-15 至 2025-10-15 共收录 9 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 9 篇

2508.00378 2025-10-15 cs.AI cs.CV 88%

CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding

Shixin Yi, Lin Shang

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);VLM(abstract);LLaVA(abstract)

Comments The paper is not yet mature and needs further improvement

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11852 2025-10-15 cs.LG 83%

Evaluating Open-Source Vision-Language Models for Multimodal Sarcasm Detection

Saroj Basnet, Shafkat Farabi, Tharindu Ranasinghe, Diptesh Kanoji, Marcos Zampieri

机构 * George Mason University(乔治·马歇尔大学) Lancaster University(兰卡斯特大学) University of Surrey(萨里大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);LLaVA(abstract);分类 cs.LG

Comments Accepted to ICDMW 2025 Workshop on Multimodal AI (MMAI). Full workshop info: https://icdmw25mmai.github.io/

Journal ref Proc. IEEE International Conference on Data Mining Workshops (ICDMW 2025), Workshop on Multimodal AI (MMAI 2025), Los Angeles, USA, December 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08307 2025-10-15 cs.CV cs.RO 83%

DSM: Constructing a Diverse Semantic Map for 3D Visual Grounding

Qinghongbing Xie, Zijian Liang, Fuhao Li, Long Zeng

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China(清华大学深圳国际研究生院,清华大学,深圳,中国)

专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract);分类 cs.CV

Comments 8 pages, 6 figures, Project Page: https://binicey.github.io/DSM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12798 2025-10-15 cs.CV 70%

Detect Anything via Next Point Prediction

Qing Jiang, Junan Huo, Xingyu Chen, Yuda Xiong, Zhaoyang Zeng, Yihao Chen, Tianhe Ren, Junzhi Yu, Lei Zhang

机构 * International Digital Economy Academy (IDEA)(国际数字经济学院(IDEA))

专题命中 视觉定位与Grounding :grounding(abstract);MLLM(abstract);分类 cs.CV

Comments homepage: https://rex-omni.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12693 2025-10-15 cs.AI 70%

ERA: Transforming VLMs into Embodied Agents via Embodied Prior Learning and Online Reinforcement Learning

Hanyang Chen, Mark Zhao, Rui Yang, Qinwei Ma, Ke Yang, Jiarui Yao, Kangrui Wang, Hao Bai, Zhenhailong Wang, Rui Pan, Mengchao Zhang, Jose Barreiros, Aykut Onol, ChengXiang Zhai, Heng Ji, Manling Li, Huan Zhang, Tong Zhang

机构 * UIUC(伊利诺伊大学香槟分校) Northwestern University(西北大学) Toyota Research Institute(丰田研究院)

专题命中 视觉定位与Grounding :vision language model(abstract);grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12763 2025-10-15 eess.SP cs.AI q-bio.QM 57%

Disentangling Neurodegeneration with Brain Age Gap Prediction Models: A Graph Signal Processing Perspective

Saurabh Sihag, Gonzalo Mateos, Alejandro Ribeiro

机构 * Department of Electrical and Computer Engineering at the University at Albany, SUNY(纽约州立大学阿尔巴尼分校电气与计算机工程系) Department of Electrical and Computer Engineering at the University of Rochester(罗切斯特大学电气与计算机工程系) Department of Electrical and Systems Engineering at the University of Pennsylvania(宾夕法尼亚大学电气与系统工程系)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted for publication in IEEE Signal Processing Magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12766 2025-10-15 cs.CL 50%

Language Models Model Language

Łukasz Borchmann

机构 * Snowflake AI Research(Snowflake人工智能研究)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12211 2025-10-15 cs.IR 50%

Reinforced Preference Optimization for Recommendation

Junfei Tan, Yuxin Chen, An Zhang, Junguang Jiang, Bin Liu, Ziru Xu, Han Zhu, Jian Xu, Bo Zheng, Xiang Wang

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11897 2025-10-15 cs.HC 50%

A Longitudinal Study on Different Annotator Feedback Loops in Complex RAG Tasks

Sara Rosenthal, Maeda Hanafi, Yannis Katsis, Lucian Popa, Marina Danilevsky

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 26 pages

详情

展开后加载摘要…

URL PDF HTML 收藏