Fine-Grained Grounding for Multimodal Speech Recognition
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments Accepted to Findings of EMNLP 2020
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments Accepted to Findings of EMNLP 2020
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments 21st IFAC World Congress, 2020, to appear
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments Best Paper Runner-up at INLG 2019 (12th International Conference on Natural Language Generation)
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI、cs.LG
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments EMNLP 2019 (5 pages + 1 references)
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments 11 pages, 5 figures, accepted by EMNLP-IJCNLP 2019
专题命中 视觉定位与Grounding :grounding(title,abstract)
专题命中 视觉定位与Grounding :grounding(title,abstract)
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments ACL 2019
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments 8 pages, 11 figures, 1 table
Journal ref IEEE Robotics and Automation Letters, vol. 4, no. 2, pp. 351-358, April 2019
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments Published in ICJAI 2017
专题命中 视觉定位与Grounding :grounding(title,abstract)
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments Paper presented at the 34nd International Conference on Logic Programming (ICLP 2018), Oxford, UK, July 14 to July 17, 2018 18 pages, LaTeX
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments Submitted to the Journal of Artificial Intelligence Research
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments 2 pages
专题命中 视觉定位与Grounding :grounding(title,abstract)
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI、cs.LG
Comments 6 pages, 4 figures, ICDL-Epirob 2016
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments 10 pages, 7 figures, 3 tables. To appear in ACL-IJCNLP 2015
专题命中 视觉定位与Grounding :grounding(title,abstract)
Journal ref Journal of Artificial Intelligence Research, feb 2015, volume 52, pages 235-286
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments 4 pages, 11 figures
Journal ref Proceedings of the 27th Symposium On Fusion Technology (SOFT-27); Liege, Belgium, September 24-28, 2012. Fusion Engineering and Design, Vol.88, Issues 9-10, October 2013, p.2100-2104
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments 13 pages, 5 figures
Journal ref International Journal of Web & Semantic Technology (IJWesT) Vol.2, No.4, 2011, 67-79
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments expanded article
专题命中 视觉定位与Grounding :VLM(title,abstract)
Comments To appear in the April 10, 2003 issue of The Astrophysical Journal 30 pages, 17 figures
Journal ref Astrophys.J. 587 (2003) 407-422
专题命中 视觉定位与Grounding :grounding(title,comments);分类 cs.CV、cs.AI
Comments 3st Place in PIC Makeup Temporal Video Grounding (MTVG) Challenge in ACM-MM 2022
UniGeo:用于文本引导跨视角地理定位的多模态大语言模型
机构 * School of Computer Engineering and Science, Shanghai University(上海大学计算机工程与科学学院) ; Institute of Collaborative Innovation, University of Macau(澳门大学协同创新研究院)
专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV
AI总结 UniGeo是一种统一多模态大语言模型,通过地理语义学习、跨视角生成及即插即用验证模块,在文本引导无人机地理定位任务中显著提升了检索性能。
从声音到症状:面向对话式医疗智能体的实时呼吸信号理解
机构 * Hippocratic AI(希波克拉底人工智能公司)
专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.AI
AI总结 提出面向对话式医疗智能体的HealthCUES系统,可实时检测分析咳嗽等呼吸信号,经内部与外部数据集评估及医疗人员验证,性能优异且具临床实用性。
Comments Accepted for publication at SIGDIAL 2026
以标注为回滚:面向视频多模态大语言模型的高效可扩展强化学习
机构 * Nankai University(南开大学) ; Northwestern Polytechnical University(西北工业大学) ; Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; NKIARI
专题命中 视觉定位与Grounding :MLLM(summary_cn);multimodal large language model(abstract);分类 cs.CV
AI总结 本文提出OraRL算法,将标注作为神谕回滚解耦优势估计,实现高效可扩展的视频MLLM强化学习,在多项视频感知基准上超越现有模型,解码速度大幅提升。
Comments Project page: https://orarl.github.io/
通过结构化奖励强化视频MLLMs的一致性
机构 * Rutgers University(罗格斯大学) ; University of Toronto(多伦多大学)
专题命中 视觉定位与Grounding :grounding(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV
AI总结 研究通过结构化奖励提升视频MLLMs的一致性,发现传统监督不足,提出结合事实和时间单元的奖励机制,提升视频理解准确性。
Comments Accepted by COLM 2026
增强视觉-语言能力的半监督医学图像分割基础模型
机构 * ECE, Northwestern University(电气工程与计算机科学系,西北大学) ; Stats, Northwestern University(统计学系,西北大学) ; Radiology, Northwestern University(放射学系,西北大学)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV
AI总结 本文提出VESSA模型,通过增强视觉-语言能力的半监督方法提升医学图像分割精度,实验表明其在有限标注条件下表现优于现有方法。
AmalthAI:面向文化遗产的开源计算机视觉平台
机构 * Democritus University of Thrace(德谟克利特色雷斯大学) ; Athena Research Center(雅典娜研究中心) ; National and Kapodistrian University of Athens(雅典国立卡波迪斯特里亚大学) ; University of Warsaw(华沙大学)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV
AI总结 AmalthAI是面向文化遗产领域专家的开源计算机视觉平台,通过集成Kubeflow、Katib等工具,支持数据集管理、模型训练与推理,可保障敏感考古数据安全,已在黏土织物印痕数据集上验证其功能。