Visual Representations inside the Language Model
机构 * University of Washington(华盛顿大学) ; University of California Los Angeles(加州大学洛杉矶分校) ; Allen Institute for AI(人工智能研究院)
专题命中 视觉定位与Grounding :LLaVA(abstract);分类 cs.CV
Comments Accepted to COLM 2025
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * University of Washington(华盛顿大学) ; University of California Los Angeles(加州大学洛杉矶分校) ; Allen Institute for AI(人工智能研究院)
专题命中 视觉定位与Grounding :LLaVA(abstract);分类 cs.CV
Comments Accepted to COLM 2025
机构 * Department of Data Science and AI, IIT Madras, India(数据科学与人工智能系,印度理工学院马德拉斯学院) ; Department of Engineering Design, IIT Madras, India(工程设计系,印度理工学院马德拉斯学院) ; LoveForm Health Technologies, India(LoveForm健康科技公司,印度) ; Department of Radiology and Imaging Sciences, Sri Ramachandra Institute of Higher Education and Research, India(放射学与成像科学系, Sri Ramachandra高等教育与研究学院,印度) ; Department of Neuro and Interventional Radiology, Sri Ramachandra Institute of Higher Education and Research, India(神经放射学与介入放射学系,Sri Ramachandra高等教育与研究学院,印度)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
Comments Paper published at "Agentic AI for Medicine" Workshop, MICCAI 2025
Journal ref Lecture Notes in Computer Science, vol 16147, 2025. Springer, Cham
机构 * Arizona State University(亚利桑那州立大学) ; Indian Institute of Technology, Kharagpur(印度理工学院,克拉格浦)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV
Comments 17 pages, 9 figures, 6 tables. Presents TimeWarp, a synthetic preference data framework to improve temporal understanding in Video-LLMs, showing consistent gains across seven benchmarks. Includes supplementary material in the Appendix
机构 * Beijing Normal–Hong Kong Baptist University(北京师范大学-香港 Baptist大学)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
机构 * Arizona State University(亚利桑那州立大学) ; University of Kansas(堪萨斯大学) ; University of Notre Dame(圣母大学) ; Duke University(杜克大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
机构 * ETH Zürich(苏黎世联邦理工学院) ; Microsoft(微软) ; TU Munich(慕尼黑工业大学) ; MCML
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV
Comments GCPR 2025 (oral presentation; Best Master's Thesis Award)
机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; Korea University(韩国大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
Comments Accepted to EMNLP 2025 Findings
机构 * Department of Biomedical Informatics, Stony Brook University, NY, USA(生物医学信息学系,石溪大学,纽约,美国) ; Department of Computer Science, Stony Brook University, NY, USA(计算机科学系,石溪大学,纽约,美国)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
Comments 16 pages, 5 figures, 1 table
专题命中 视觉定位与Grounding :grounding(abstract)
Comments 34 pages including appendices; figures included. Primary subject class: q-fin.TR. Cross-lists: cs.LG; q-fin.CP
专题命中 视觉定位与Grounding :grounding(abstract)
Comments 12 pages, 2 figures, ICFP '25 The miniKanren and Relational Programming Workshop
机构 * Carnegie Mellon University(卡内基梅隆大学)
专题命中 视觉定位与Grounding :grounding(abstract)
Comments Accepted at the NeurIPS 2025 Workshop on Space in Vision, Language, and Embodied AI (SpaVLE). *Equal contribution
机构 * MIPT(莫斯科国立交通大学) ; Sberbank of Russia, Robotics Center(俄罗斯储蓄银行机器人中心) ; AIRI
专题命中 视觉定位与Grounding :visual language model(abstract)
Comments Accepted to IROS 2025
机构 * DaVInCi Laboratory(DaVInCi实验室) ; University of Cincinnati(克利夫兰大学)
专题命中 文档图表理解 :VLM(abstract);分类 cs.AI
Journal ref The 2nd MERCADO Workshop at IEEE VIS 2025: Multimodal Experiences for Remote Communication Around Data Online, IEEE VIS 2025
专题命中 文档图表理解 :vision language model(abstract);分类 cs.LG
Comments EMNLP 2025
机构 * Opus AI Research(Opus人工智能研究机构) ; Brown University(布朗大学) ; Zhejiang University(浙江大学) ; University of Toronto(多伦多大学)
专题命中 文档图表理解 :MLLM(abstract);分类 cs.CV
Comments working in progress
专题命中 GUI与屏幕智能体 :MLLM(title,abstract)
机构 * KAIST(韩国科学技术院) ; RLWRLD ; UC Berkeley(伯克利大学)
专题命中 GUI与屏幕智能体 :vision-language model(abstract);VLM(abstract);分类 cs.AI
Comments Project page: https://huiwon-jang.github.io/contextvla
机构 * Stony Brook University(石溪大学) ; AWS AI Labs(亚马逊人工智能实验室)
专题命中 GUI与屏幕智能体 :vision-language model(abstract);VLM(abstract);分类 cs.AI
机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学全球化人工智能学院) ; Department of Data Science, City University of Hong Kong(香港城市大学数据科学系) ; Alibaba Group(阿里巴巴集团)
专题命中 GUI与屏幕智能体 :multimodal large language model(abstract);分类 cs.CV
Comments Under Review
专题命中 GUI与屏幕智能体 :grounding(abstract);分类 cs.AI
机构 * VLM Safety LAB, MODULABS(视觉语言模型安全实验室,MODULABS) ; ETRI(电子技术研究院) ; KAIST(韩国科学技术院)
专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV、cs.AI
Comments Accepted to Safe and Trustworthy Multimodal AI Systems(SafeMM-AI) Workshop at ICCV2025, Non-archival track
机构 * Université de Montréal(蒙特利尔大学) ; Mila – Quebec AI Institute(魁北克人工智能研究所)
专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV、cs.AI
机构 * The University of Queensland(昆士兰大学) ; Nanjing University(南京大学) ; University of California, Los Angeles(加州大学洛杉矶分校) ; University of California, Merced(加州大学默塞德分校)
专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.AI
专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV
Comments 33 pages, 11 figures, 7 tables
机构 * The Hong Kong Polytechnic University(香港理工大学) ; Tsinghua University(清华大学) ; InspireOmni AI ; Alibaba Group(阿里巴巴集团) ; Case Western Reserve University(凯斯西储大学)
专题命中 VLM训练与架构 :VLM(title,abstract);vision language model(abstract);分类 cs.CV
Comments Accepted at NeurIPS 2025
机构 * Stanford University(斯坦福大学) ; NVIDIA USA(NVIDIA公司)
专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV
机构 * Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) ; Large Language Model Department, Tencent(腾讯大语言模型部门) ; School of Computer and Communication Engineering, University of Science and Technology Beijing(北京科技大学计算机与通信工程学院) ; Faculty of Science and Technology, University of Macau(澳门大学科技学院)
专题命中 VLM训练与架构 :vision-language model(title);visual language model(abstract);分类 cs.AI
Comments Accepted by EMNLP 2025 findings
机构 * Google DeepMind(谷歌DeepMind) ; University of Central Florida(中央佛罗里达大学)
专题命中 VLM训练与架构 :vision-language model(title);分类 cs.CV、cs.AI
机构 * NAVER AI Lab(NAVER AI实验室)
专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.LG
Comments Code: https://github.com/naver-ai/prolip HuggingFace Hub: https://huggingface.co/collections/SanghyukChun/prolip-6712595dfc87fd8597350291 33 pages, 4.5 MB; LongProLIP paper: arXiv:2503.08048; Multiplicity paper for more background: arxiv.org:2505.19614; v4: fix typos