IMAGINATOR: Pre-Trained Image+Text Joint Embeddings using Word-Level Grounding of Images
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI
专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV、cs.AI
Comments Accepted in Vision Datasets Understanding at CVPR 2023
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.LG
Comments ECCV 2022; 25 pages (including Supplementary Materials); Updated related works
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG
Comments Accepted to AAAI 2022. Code at https://github.com/styfeng/VisCTG
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG
Comments Corrected references
Journal ref 1st Workshop on Human and Machine Decisions (WHMD 2021), NeurIPS 2021
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.LG
Comments accepted at CIFMA2021 (https://cifma.github.io/)
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.LG
Journal ref Conference on Computer Vision and Pattern Recognition. 2020, pp. 2220-2229
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.LG
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG
Comments JAIR 2018
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG
Comments 23 pages, 7 figures, published in Neurocomputing
Journal ref Neurocomputing, Volume 268, 13 December 2017, Pages 142-152
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG
Comments 8 pages, 4 figures, workshop at IROS 2015 conference
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG
Comments 59 pages, 16 figures, submitted to Neural Networks
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG
Comments 6 pages, 3 figures
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.LG
Comments Under review on Journal of Artificial Intelligence Research (JAIR) -- Special Track on Deep Learning, Knowledge Representation, and Reasoning
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG
Comments 14 pages, 4 figures
专题命中 视觉定位与Grounding :grounding(title,comments);分类 cs.LG
Comments D. Matthews, S. Kriegman, C. Cappelle and J. Bongard, "Word2vec to behavior: morphology facilitates the grounding of language in machines," 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Macau, China, 2019. \c{opyright} 2019 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses
专题命中 视觉定位与Grounding :grounding(title,comments);分类 cs.AI
Comments 9 pages, 8 figures, To appear in the Proceedings of the ACL workshop Language Grounding for Robotics, Vancouver, Canada
HODAgent:面向物理世界人机交互的按需、响应式人形机器人
机构 * Xiaopeng(小鹏)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);grounding(abstract)
AI总结 HODAgent作为面向服务人形机器人的System-2具身智能体,通过半双工架构实现自适应服务,在仿真与实体机器人上均显著优于基准模型。
Comments we have received a formal directive from our company requiring all company assets to undergo a mandatory internal review process before any public release. We are now required to immediately withdraw the paper to comply with this policy
你会点击什么?基于偏好感知高光检索的个性化视频缩略图生成
专题命中 视觉定位与Grounding :VLM(summary_cn,abstract)
AI总结 研究针对视频平台个性化视频缩略图生成问题,提出两阶段框架,先通过偏好感知检索选取视觉锚点,再经VLM引导的扩散管道生成缩略图,实验和用户研究证明该方法性能优且能提升用户参与度。
ProAct:用于结构感知主动响应的基准和多模态框架
机构 * Department of Computer Science and Engineering, The Hong Kong University of Science and Technology (HKUST), Hong Kong SAR, China(香港科技大学计算机科学与工程系) ; Tencent, Shenzhen, China(腾讯(中国深圳)) ; Shenzhen Institute of Advanced Technology (SIAT), Chinese Academy of Sciences, Shenzhen, China(深圳先进技术研究所(SIAT),中国科学院)
专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract)
AI总结 针对主动智能体开发受资源缺乏阻碍的问题,引入ProAct-75基准,提出由多模态大语言模型驱动的ProAct-Helper,其利用任务图进行行动选择,实验证明该方法在触发检测等方面优于闭源模型。
图像无法表达的:语言引导的嗅觉表征学习
机构 * LIX, École Polytechnique, IP Paris, CNRS(LIX,巴黎高等理工学院,IP巴黎,CNRS)
专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 研究图像与嗅觉对齐难题,提出SCENT多模态框架,利用视觉语言模型生成场景描述符,训练气味编码器,通过语言引导潜在分解分离特定对象气味与环境贡献,提升跨模态检索性能,产生可解释表征。
Comments ECCV 2026. Project page: https://www.lix.polytechnique.fr/vista/projects/2026_scent_tsonis/
GAVEL:基于视觉定位的标题错误验证与定位
机构 * OMRON SINIC X Corporation(欧姆龙SINIC X公司)
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract_cn);grounding(abstract)
AI总结 提出GAVEL任务,联合解决图像-文本对的验证、解释和定位问题,构建数据集和基准,监督基线在定位和解释指标上持续改进。
Comments conference
面向机器的对称熵约束视频编码
专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);grounding(abstract)
AI总结 提出SEC-VCM框架,通过双向熵约束机制对称对齐视频编解码器与视觉骨干网络,保留语义并丢弃无关信息,结合语义-像素双路径融合提升机器视觉任务性能,在多项任务上显著优于H.266/VVC。
Comments Accepted by IEEE Transactions on Image Processing. This is the author's accepted manuscript (AAM)
视频语言模型何时停止观看?多模态RLVR中视觉捷径的形成与逆转受奖励强度控制
机构 * Zekun Xu(徐泽坤)
专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 研究多模态强化学习中视觉捷径的形成与逆转机制,发现惩罚强度可控制其出现和消除,且存在关键干预窗口。
Comments 11 pages, 4 figures
Dementia-Agents:用于痴呆分期和表型分析的多模态多智能体系统
机构 * Monash University(莫纳什大学) ; Eastern Health(东部健康) ; Lived Experience Advisor(生活经验顾问)
专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract_cn)
AI总结 提出Dementia-Agents多智能体框架,通过数据翻译、专家预测和协调聚合三步流程,在真实临床队列中实现优于单一MLLM和先前多智能体系统的痴呆综合征级分期与表型分析。
Comments 8 pages
一图足矣:基于文本世界模型的智能体单样本图像生成用于长尾空间感知
机构 * Tsinghua University(清华大学) ; SenseTime Research(商汤科技研究院) ; Sun Yat-Sen University(中山大学) ; Technical University of Munich(慕尼黑工业大学) ; Heilbronn Data Science Center(海尔布隆数据科学中心) ; Munich Data Institute(慕尼黑数据研究所)
专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 提出WMGen-v1框架,利用单张参考图像通过LVLM构建结构化场景表示,LLM进行物理合理的场景扩展,再由扩散模型生成多样化的长尾训练数据,缓解空间感知中的数据稀缺问题。
MALLVI:一种多智能体框架用于集成通用机器人操作
机构 * Department of Electrical Engineering, Sharif University of Technology(电气工程系,谢里夫大学)
专题命中 视觉定位与Grounding :vision language model(abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 MALLVI通过多智能体协作实现闭环反馈驱动的机器人操作,提升泛化能力和零样本任务成功率。
Comments Some fundemental change in text and codebase
CURE:基于课程引导的多任务训练实现可靠的解剖学接地报告生成
机构 * Pontificia Universidad Católica de Chile(智利天主教大学) ; CENIA ; iHEALTH ; KAUST(科威特皇家科学与技术局)
专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 提出CURE框架,通过课程学习动态调整多任务训练,提升医学报告生成的视觉接地准确性和事实一致性,无需额外数据。
Comments 31 pages, 7 figures, accepted to CVPR 2026 (oral)
Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026, pp. 36279-36289
VAMPS: 视觉辅助数学问题求解基准
机构 * University of California, Berkeley(加州大学伯克利分校)
专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 提出VAMPS基准,通过1,168道双语多选题评估多模态大模型在借助绘图工具进行数学推理时的表现,发现直接解析求解优于工具辅助视觉求解。
FOCUS: 通过视觉支持约束和策略优化强制上下文目标定位
机构 * Amazon, Seattle, USA(亚马逊(美国西雅图))
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract_cn);分类 cs.CV、cs.AI、cs.LG
AI总结 提出一种两阶段训练框架,通过优化支持框与查询图像间的上下文注意力并结合GRPO强化学习,实现无类别监督的类别无关上下文目标定位,7B模型性能超越72B模型。
Comments Accepted at ICML 2026. * Equal Contributions