arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-12-23 至 2025-12-23 共收录 18 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 18 篇

2402.06118 2025-12-23 cs.CV cs.AI 90%

ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine-Grained Reward Modeling

ViGoR:通过细粒度奖励建模提升大视觉语言模型的视觉 grounding

Siming Yan, Min Bai, Weifeng Chen, Xiong Zhou, Qixing Huang, Li Erran Li

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) AWS AI(AWS人工智能)

专题命中 视觉定位与Grounding :vision language model(title,abstract);grounding(title,abstract);visual reasoning(abstract);分类 cs.CV、cs.AI

AI总结 ViGoR通过细粒度奖励建模提升大视觉语言模型的视觉 grounding 能力,采用更经济的人类评估和自动化方法,有效提高视觉推理准确性。

Comments Accepted by ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18878 2025-12-23 cs.CV cs.AI 86%

CrashChat: A Multimodal Large Language Model for Multitask Traffic Crash Video Analysis

CrashChat: 一种多模态大语言模型用于多任务交通事故视频分析

Kaidi Liang, Ke Li, Xianbiao Hu, Ruwen Qin

机构 * Stony Brook University(石溪大学) The Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 CrashChat是一种多模态大语言模型,用于多任务交通事故视频分析,通过任务解耦和分组策略提升事故识别、定位等任务的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18947 2025-12-23 cs.CV 85%

OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language Model

OpenHOI: 基于多模态大语言模型的开放世界手-物体交互合成

Zhenhao Zhang, Ye Shi, Lingxiao Yang, Suting Ni, Qi Ye, Jingya Wang

机构 * ShanghaiTech University(上海科技大学) Zhejiang University(浙江大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);MLLM(abstract);分类 cs.CV

AI总结 OpenHOI通过多模态大语言模型实现开放世界手-物体交互合成,能生成长周期操控序列并处理复杂语言指令。

Comments Accepted by NeurIPS 2025 as Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12715 2025-12-23 cs.CV cs.RO 83%

AsyMoE: Leveraging Modal Asymmetry for Enhanced Expert Specialization in Large Vision-Language Models

AsyMoE:利用模态不对称性提升大视觉-语言模型专家专业化

Heng Zhang, Haichuan Hu, Yaomin Shen, Weihao Yu, Yilei Yuan, Haochen You, Guo Cheng, Zijian Zhang, Lubin Gan, Huihui Wei, Hao Zhang, Jin Huang

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV

AI总结 AsyMoE通过三个专门专家组解决视觉-语言模态不对称问题,提升大模型专家专业化性能。

Comments This submission has been withdrawn by the authors due to a fundamental error in the methodology that affects the validity of the main results

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21728 2025-12-23 cs.CL cs.AI cs.LG 81%

Affective Multimodal Agents with Proactive Knowledge Grounding for Emotionally Aligned Marketing Dialogue

具有主动知识 grounding 的情感多模态代理用于情感对齐的营销对话

Lin Yu, Xiaofei Han, Yifei Kang, Chiung-Yi Tseng, Danyang Zhang, Ziqian Bi, Zhimo Han

机构 * Hunan Police Academy, Department of Criminal Investigation(湖南警察学院犯罪侦查系) Business College, California State University(加州州立大学长滩分校商学院) Northwestern University(西北大学) AI Agent Lab, Vokram Group(Vokram集团AI代理实验室) Beijing University of Technology(北京理工大学) Zheng Zhou University of Light Industry(郑州轻工业大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI、cs.LG

AI总结 AffectMind通过主动知识 grounding 和情绪-意图对齐模型,提升营销对话中的情感一致性与说服效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20732 2025-12-23 cs.MM cs.CV 79%

Prompt-Aware Adaptive Elastic Weight Consolidation for Continual Learning in Medical Vision-Language Models

面向提示的自适应弹性权重固化用于医学视觉-语言模型的持续学习

Ziyuan Gao, Philippe Morel

机构 * University College London(伦敦大学学院)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

AI总结 PA-EWC通过提示引导的参数专门化,有效缓解医疗视觉-语言模型在持续学习中的灾难性遗忘问题,提升多模态医学任务的性能。

Comments Accepted by 32nd International Conference on MultiMedia Modeling (MMM 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11792 2025-12-23 cs.AI 74%

Solver-Informed RL: Grounding Large Language Models for Authentic Optimization Modeling

Solver-Informed RL: 为真实优化建模奠定大语言模型基础

Yitian Chen, Jingfan Xia, Siyu Shao, Dongdong Ge, Yinyu Ye

机构 * Cardinal Operations, China(中国卡迪纳尔运营公司) Shanghai University of Finance and Economics(上海财经大学) The University of Hong Kong(香港大学) Antai School of Economics and Management, Shanghai Jiao Tong University(上海交通大学安泰经济管理学院) Department of Management Science and Engineering, Stanford University(斯坦福大学管理科学与工程系)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

AI总结 SIRL通过强化学习与外部优化求解器结合,提升大语言模型在优化建模中的准确性与实用性。

Journal ref 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19678 2025-12-23 cs.CV cs.AI 62%

WorldWarp: Propagating 3D Geometry with Asynchronous Video Diffusion

WorldWarp: 通过异步视频扩散传播3D几何

Hanyang Kong, Xingyi Yang, Xiaoxu Zheng, Xinchao Wang

机构 * National University of Singapore(新加坡国立大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

AI总结 WorldWarp通过结合3D结构锚和2D生成细化器,利用时空扩散模型实现几何一致的视频生成,解决遮挡和复杂轨迹问题。

Comments Project page: https://hyokong.github.io/worldwarp-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18986 2025-12-23 cs.LG cs.AI 62%

R-GenIMA: Integrating Neuroimaging and Genetics with Interpretable Multimodal AI for Alzheimer's Disease Progression

R-GenIMA:整合神经影像与基因组学的可解释多模态AI用于阿尔茨海默病进展

Kun Zhao, Siyuan Dai, Yingying Zhang, Guodong Liu, Pengfei Gu, Chenghua Lin, Paul M. Thompson, Alex Leow, Heng Huang, Lifang He, Liang Zhan, Haoteng Tang

机构 * Eli and Lilly company(艾利和利公司)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.AI、cs.LG

AI总结 R-GenIMA通过整合神经影像与基因组学,利用可解释的多模态AI方法,实现了对阿尔茨海默病进展的精准预测与机制揭示。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18177 2025-12-23 cs.AI cs.CV 62%

NEURO-GUARD: Neuro-Symbolic Generalization and Unbiased Adaptive Routing for Diagnostics -- Explainable Medical AI

NEURO-GUARD:神经符号泛化与无偏自适应路由用于诊断——可解释的医疗AI

Midhat Urooj, Ayan Banerjee, Sandeep Gupta

机构 * Arizona State University(亚利桑那州立大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

AI总结 NEURO-GUARD通过整合视觉变换器与语言驱动推理,提升医疗图像诊断的准确性、透明性和泛化能力。

Comments Accepted at Asilomar Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00908 2025-12-23 cs.CV cs.GR 57%

GraphGeo: Multi-Agent Debate Framework for Visual Geo-localization with Heterogeneous Graph Neural Networks

GraphGeo: 多智能体辩论框架用于具有异构图神经网络的视觉地定位

Heng Zheng, Yuling Shi, Xiaodong Gu, Haochen You, Zijian Zhang, Lubin Gan, Hao Zhang, Wenjun Huang, Jin Huang

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 GraphGeo通过异构图神经网络的多智能体辩论框架,提升视觉地定位的准确性与鲁棒性。

Comments This submission has been withdrawn by the authors due to a fundamental error in the methodology that affects the validity of the main results

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18853 2025-12-23 cs.CV cs.HC 57%

VizDefender: Unmasking Visualization Tampering through Proactive Localization and Intent Inference

VizDefender:通过主动定位和意图推断揭示可视化篡改

Sicheng Song, Yanjie Zhang, Zixin Chen, Huamin Qu, Changbo Wang, Chenhui Li

机构 * East China Normal University(华东师范大学) Hong Kong University of Science and Technology(香港科学与技术大学) School of Computer Science and Technology(计算机科学与技术学院)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

AI总结 VizDefender通过半脆弱水印和意图分析模块,主动定位和推断可视化篡改,有效检测和分析数据篡改与视觉编码篡改。

Comments IEEE Transactions on Visualization and Computer Graphics (IEEE PacificVis'26 TVCG Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17642 2025-12-23 cs.CV 57%

Efficient Redundancy Reduction for Open-Vocabulary Semantic Segmentation

高效开放词汇语义分割中的冗余减少

Lin Chen, Qi Yang, Kun Ding, Zhihao Li, Gang Shen, Fei Li, Qiyuan Cao, Shiming Xiang

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室(MAIS)、自动化研究所、中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) School of Software, Shandong University(山东大学软件学院) China Tower Corporation Limited(中国铁塔股份有限公司)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 本文提出ERR-Seg,通过减少冗余信息和优化序列建模,提升开放词汇语义分割的效率与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18440 2025-12-23 cs.CL cs.AI 57%

An Agentic AI Framework for Training General Practitioner Student Skills

一种用于培训全科医学生技能的代理AI框架

Victor De Marez, Jens Van Nooten, Luna De Bruyne, Walter Daelemans

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出了一种代理AI框架,用于培训全科医学生技能,通过统一病例生成、角色驱动对话和基于标准的评估,提升医学教育中虚拟模拟患者的效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19466 2025-12-23 cs.CY cs.CL cs.HC 50%

Epistemological Fault Lines Between Human and Artificial Intelligence

人类与人工智能之间的知识论断层

Walter Quattrociocchi, Valerio Capraro, Matjaž Perc

机构 * Department of Computer Science, Sapienza University of Rome, Rome, Italy Department of Psychology, University of Milan Bicocca, Milan, Italy Faculty of Natural Sciences Mathematics, University of Maribor, Maribor, Slovenia Community Healthcare Center Dr. Adolf Drolc Maribor, Maribor, Slovenia University College, Korea University, Seoul, Republic of Korea Department of Physics, Kyung Hee University, Seoul, Republic of Korea

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文揭示大型语言模型与人类认知在知识生成机制上的结构性差异,指出LLMs是随机模式完成系统而非知识代理,并识别七种知识断层,对社会评估、治理及知识素养提出影响。

Comments 16 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22876 2025-12-23 cs.MA 50%

Cooperation as a Black Box: Conceptual Fluctuation and Diagnostic Tools for Misalignment in MAS

协作作为黑箱:多智能体系统中概念波动与对齐偏差的诊断工具

Shayak Nandi, Fernanda M. Eliott

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文提出马赛克框架,用于诊断多智能体系统中概念波动与对齐偏差,强调术语一致性与道德基础以确保系统技术与道德对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18225 2025-12-23 cs.CL cs.SI 50%

GeoSense-AI: Fast Location Inference from Crisis Microblogs

GeoSense-AI:从危机推文快速推断位置

Deepit Sapru

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 GeoSense-AI通过领域调优的NLP和知识基础,从危机推文快速推断位置,提升应急响应效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17992 2025-12-23 cs.RO 50%

Unifying Deep Predicate Invention with Pre-trained Foundation Models

统一深度谓词发明与预训练基础模型

Qianwei Wang, Bowen Li, Zhanpeng Luo, Yifan Xu, Alexander Gray, Tom Silver, Sebastian Scherer, Katia Sycara, Yaqi Xie

机构 * Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所) Computer Science and Engineering Division, University of Michigan(密歇根大学计算机科学与工程系) Department of Computer Science, University of Pittsburgh(匹兹堡大学计算机科学系) Centaur AI Institute(Centaur人工智能研究所) Department of Electrical and Computer Engineering, Princeton University(普林斯顿大学电气与计算机工程系)

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 UniPred通过双层学习框架统一深度谓词发明与预训练基础模型,提升机器人任务的可扩展性和灵活性。

Comments 18 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏