arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-11-07 至 2025-11-07 共收录 6 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 6 篇

2511.03757 2025-11-07 cs.LG cs.AI 73%

Laugh, Relate, Engage: Stylized Comment Generation for Short Videos

Xuan Ouyang, Senan Wang, Bouzhou Wang, Siyuan Xiahou, Jinrong Zhou, Yuekang Li

机构 * University of New South Wales(新南威尔士大学) University of Sydney(悉尼大学) The University of Hong Kong(香港大学) University of Southern California(南加州大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01019 2025-11-07 cs.CL cs.AI cs.CE cs.LG physics.ao-ph 62%

OceanAI: A Conversational Platform for Accurate, Transparent, Near-Real-Time Oceanographic Insights

Bowen Chen, Jayesh Gajbhar, Gregory Dusek, Rob Redmon, Patrick Hogan, Paul Liu, DelWayne Bohnenstiehl, Dongkuan Xu, Ruoying He

机构 * North Carolina State University(北卡罗来纳州立大学) NOAA(国家海洋和大气管理局)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments A related presentation will be given at the AGU(American Geophysical Union) and AMS(American Meteorological Society) Annual Meetings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04590 2025-11-07 cs.LG cs.IT math.IT 57%

Complexity as Advantage: A Regret-Based Perspective on Emergent Structure

Oshri Naparstek

机构 * IBM Research(IBM研究院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments 15 pages. Under preparation for submission to ICML 2026. Feedback welcome

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23769 2025-11-07 cs.CV 57%

TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models

Yao Xiao, Qiqian Fu, Heyi Tao, Yuqun Wu, Zhen Zhu, Derek Hoiem

机构 * Siebel School of Computing and Data Science(塞比尔计算与数据科学学院) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Published in TMLR, with a J2C Certification

Journal ref Transactions on Machine Learning Research, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04847 2025-11-07 cs.CL cs.AI 57%

Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards

Manveer Singh Tamber, Forrest Sheng Bao, Chenyu Xu, Ge Luo, Suleman Kazi, Minseok Bae, Miaoran Li, Ofer Mendelevitch, Renyi Qu, Jimmy Lin

机构 * University of Waterloo(滑铁卢大学) Vectara(Vectara公司) Iowa State University(爱荷华州立大学) Stanford University(斯坦福大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments EMNLP Industry Track 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04112 2025-11-07 cs.CV 57%

SpatialLock: Precise Spatial Control in Text-to-Image Synthesis

Biao Liu, Yuanzhi Liang

机构 * The Sugon Group(神舟集团) TeleAI China Telecom(中国电信) Shanghai China(上海中国)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏