Mirror Duality in a Spencer-Type Complex: Analytic and Riemann-Roch Perspectives
专题命中 视觉定位与Grounding :grounding(abstract)
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉定位与Grounding :grounding(abstract)
机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Peking University(北京大学) ; Shanghai Jiao Tong University(上海交通大学)
专题命中 文档图表理解 :vision-language model(title,abstract);分类 cs.CV
Comments Technical Report; GitHub Repo: https://github.com/opendatalab/MinerU Hugging Face Model: https://huggingface.co/opendatalab/MinerU2.5-2509-1.2B Hugging Face Demo: https://huggingface.co/spaces/opendatalab/MinerU
机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; Alibaba Cloud Computing(阿里云计算)
专题命中 文档图表理解 :vision-language model(abstract)
Comments Under review
机构 * Qualcomm AI Research(高通人工智能研究)
专题命中 GUI与屏幕智能体 :VLM(title,abstract);vision-language model(abstract);分类 cs.LG
Comments 19 pages including references, 6 figures. Accepted to CoRL LEAP 2025
机构 * Psychiatry and Biobehavioral Sciences, UCLA(乌尔拉克大学精神病学与生物行为科学系) ; Brain and Creativity Institute, USC(美国大学脑与创造力研究所) ; Google DeepMind, Zurich(谷歌深度Mind瑞士分公司)
专题命中 GUI与屏幕智能体 :multimodal large language model(title);分类 cs.CV、cs.AI
机构 * Beijing Institute of Technology(北京理工大学) ; State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI) ; DataCanvas ; Beijing University of Posts and Telecommunications(北京邮电大学) ; Shenzhen MSU-BIT University(深圳MSU-BIT大学)
专题命中 GUI与屏幕智能体 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG
机构 * Emory University(埃默里大学) ; Tencent AI Lab(腾讯AI实验室)
专题命中 GUI与屏幕智能体 :vision-language model(abstract);VLM(abstract);分类 cs.AI、cs.LG
Comments Accepted to EMNLP 2025 (Main Conference)
机构 * Apple(苹果公司)
专题命中 GUI与屏幕智能体 :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI
机构 * Huazhong University of Science and Technology(华中科技大学) ; Xiaomi EV(小米电动车)
专题命中 GUI与屏幕智能体 :vision-language model(abstract);VLM(abstract);分类 cs.CV
机构 * Centre for Artificial Intelligence, UCL(人工智能中心,伦敦大学学院) ; Qualcomm AI Research(高通人工智能研究)
专题命中 GUI与屏幕智能体 :VLM(abstract);分类 cs.CV、cs.AI、cs.LG
Comments Presented at 9th Conference on Robot Learning (CoRL 2025), Seoul, Korea
机构 * Carnegie Mellon University(卡内基梅隆大学) ; Davidson College(戴维森学院)
专题命中 GUI与屏幕智能体 :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG
机构 * Manipal University Jaipur(贾浦尔曼普尔大学) ; UNC–Charlotte(北卡罗来纳州立大学夏洛特分校) ; IIIT Hyderabad(海得拉巴印度理工学院)
专题命中 GUI与屏幕智能体 :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG
Comments 6 pages, 2 tables
机构 * Panasonic Connect Co., Ltd.(松下电器(株式会社)) ; Panasonic R&D Center(松下研发中心) ; National University of Singapore(新加坡国立大学)
专题命中 GUI与屏幕智能体 :vision-language model(abstract);分类 cs.AI、cs.LG
机构 * Intelligent Robotics Group at the Department of Electrical Engineering and Automation, School of Electrical Engineering, Aalto University(Aalto大学电气工程学院电气工程与自动化系智能机器人组) ; Biomimetics and Intelligent Systems Group at the Faculty of Information Technology and Electrical Engineering, University of Oulu(奥卢大学信息科技与电气工程学院仿生学与智能系统组) ; Section of Mechanical Technology at the Department of Engineering Technology and Didactics, Technical University of Denmark(丹麦技术大学工程技术与教学系机械技术部门)
专题命中 GUI与屏幕智能体 :vision language model(abstract)
Comments Under review for ICRA 2026
机构 * University of Science and Technology of China(中国科学技术大学) ; Microsoft Research(微软研究院) ; Nanjing University(南京大学) ; Central South University(中南大学) ; Zhejiang University(浙江大学) ; Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)
专题命中 GUI与屏幕智能体 :grounding(abstract)
机构 * Queen Mary University of London(伦敦女王学院) ; University College London(伦敦大学学院)
专题命中 GUI与屏幕智能体 :vision-language model(abstract)
机构 * AIM Intelligence(AIM智能公司) ; Yonsei University(延世大学) ; Lablup Inc.(Lablup公司)
专题命中 幻觉与鲁棒性 :vision language model(title);vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI
机构 * Savitribai Phule Pune University (SPPU)(萨维特里·布尔大学(SPPU))
专题命中 幻觉与鲁棒性 :vision-language model(title);vision language model(abstract);LLaVA(abstract);分类 cs.CV
Comments 10 pages, 12 figures. Code for MATS released at https://github.com/thubZ09/mats-spatial-reasoning
机构 * Fudan University(复旦大学) ; University of Southern California(南加州大学) ; ByteDance(字节跳动)
专题命中 幻觉与鲁棒性 :VLM(title,abstract);vision-language model(abstract);分类 cs.CV
机构 * The University of Hong Kong(香港大学)
专题命中 幻觉与鲁棒性 :vision-language model(title);vision language model(abstract);分类 cs.CV、cs.AI
机构 * Oracle AI
专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Comments Accepted in EMNLP 2025
机构 * Mohamed bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学) ; Hong Kong Baptist University(香港 Baptist大学)
专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
专题命中 幻觉与鲁棒性 :vision-language model(abstract);VLM(abstract)
机构 * Yale University(耶鲁大学)
专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.AI
机构 * Korea University(韩国大学) ; Meta GenAI(Meta 生成人工智能) ; KAIST(韩国科学技术院)
专题命中 幻觉与鲁棒性 :MLLM(abstract);分类 cs.CV
Comments ICCV 2025 Highlight
机构 * School of Computing, National University of Singapore(计算学院,新加坡国立大学)
专题命中 VLM训练与架构 :VLM(title,abstract);vision-language model(abstract);分类 cs.CV
机构 * College of Computer Science and Artificial Intelligence, Fudan University(计算机科学与人工智能学院,复旦大学)
专题命中 VLM训练与架构 :vision-language model(title,abstract);分类 cs.CV
机构 * School of Engineering, Westlake University, Hangzhou, China(西湖大学工程学院) ; School of Cyberspace Security, Nanjing University of Science and Technology, Nanjing, China(南京理工大学网络安全学院)
专题命中 VLM训练与架构 :LLaVA(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
机构 * Jadavpur University(贾瓦帕尔大学)
专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG
Comments Accepted at the IEEE/CVF International Conference on Computer Vision (ICCV 2025), Workshop on Curated Data for Efficient Learning
机构 * Pohang University of Science and Technology (POSTECH)(釜山科学技术大学) ; Tübingen AI Center(图宾根人工智能中心) ; Universität Tübingen(图宾根大学)
专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI
Comments This paper was first submitted to NeurIPS 2024 in May 2024