MindFlow: Revolutionizing E-commerce Customer Support with Multimodal LLM Agents
机构 * Xiaoduo AI(小多AI) ; University of Dayton(代顿大学)
专题命中 视觉推理 :MLLM(abstract);分类 cs.AI
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * Xiaoduo AI(小多AI) ; University of Dayton(代顿大学)
专题命中 视觉推理 :MLLM(abstract);分类 cs.AI
机构 * ECE, Seoul National University, Korea.(电子工程系,首尔国立大学,韩国) ; IPAI, Seoul National University, Korea.(人工智能研究所,首尔国立大学,韩国)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
Comments Accepted to the 42nd International Conference on Machine Learning (ICML 2025)
机构 * Fraunhofer Institute for Integrated Circuits IIS(弗劳恩霍夫集成电路研究所)
专题命中 视觉推理 :LLaVA(abstract);分类 cs.AI
Journal ref IEEE Wireless Communications and Networking Conference (WCNC), March 2025, Milan, Italy
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
Comments Technical Report: https://github.com/Kwai-Keye/Keye
专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI
机构 * School of Computing and Augmented Intelligence, Arizona State University(计算与增强智能学院,亚利桑那州立大学)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI
机构 * HKUST(香港理工大学) ; Dartmouth College(达特茅斯学院)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
Comments Project page: https://danielshkao.github.io/thinkfirst.html
机构 * Department of Aerospace Engineering and Coordinated Science Laboratory, University of Illinois at Urbana-Champaign(航空航天工程系和协调科学实验室,伊利诺伊大学厄巴纳-香槟分校) ; Department of Electrical Engineering and Coordinated Science Laboratory, University of Illinois at Urbana-Champaign(电气工程系和协调科学实验室,伊利诺伊大学厄巴纳-香槟分校) ; Intelligent Robotics Group, NASA Ames Research Center(智能机器人组,美国国家航空航天局阿姆斯研究中心)
专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV
Comments 4 pages, 5 figures, presented at the Workshop on 3D Visual Representations for Manipulation at the 2023 IEEE International Conference on Robotics and Automation in Yokohama, Japan. Video presentation [https://youtu.be/mg30uCUtpOk]. Poster [https://hollydinkel.github.io/assets/pdf/ICRA20243DVRM_poster.pdf] 3DVRM Workshop [https://3d-manipulation-workshop.github.io/]
机构 * Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
机构 * LMU Munich(慕尼黑大学) ; Munich Center for Machine Learning(慕尼黑机器学习中心) ; University of Oxford(牛津大学)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
Comments preprint version; 23 pages (including references and appendix)
机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(中国教育部多媒体可信感知与高效计算重点实验室,厦门大学) ; The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; The Chinese University of Hong Kong(香港中文大学) ; Bytedance(字节跳动) ; National University of Singapore(新加坡国立大学) ; Tsinghua University(清华大学)
专题命中 视觉推理 :MLLM(abstract);分类 cs.CV
Comments 40 pages, 26 figures
机构 * The Hong Kong Polytechnic University(香港理工大学) ; Zhejiang University(浙江大学) ; University of Electronic Science and Technology of China(电子科技大学) ; Reallm Labs(Reallm 实验室) ; Amazon(亚马逊) ; The Hong Kong University of Science and Technology(香港理工大学)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI
机构 * MIT(麻省理工学院) ; Harvard University(哈佛大学) ; University of Cambridge(剑桥大学) ; Brown University(布朗大学) ; Yale University(耶鲁大学) ; Stanford University(斯坦福大学)
专题命中 视觉推理 :VLM(abstract);分类 cs.AI
Comments 5 figures, 19 pages
机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学光华学院人工智能学院) ; Beijing Academy of Artificial Intelligence(北京人工智能研究院) ; University of Trento(特伦托大学) ; Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
机构 * Department of Biomedical Systems Informatics, Yonsei University College of Medicine(生物医学系统信息学系,延世大学医学院) ; Yonsei University College of Medicine(延世大学医学院) ; Department of Radiation Oncology, Yonsei University College of Medicine(放射肿瘤学系,延世大学医学院) ; Department of Biomedical Systems Informatics and the Department of Psychiatry, Yonsei University College of Medicine(生物医学系统信息学系和精神病学系,延世大学医学院) ; Institute for Innovation in Digital Healthcare, Yonsei University(数字医疗创新研究所,延世大学)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI
Comments 11 pages, 6 figures
机构 * The University of Hong Kong(香港大学) ; Tsinghua University(清华大学) ; University College London(伦敦大学学院)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI
机构 * School of Computing, Dublin City University(都柏林城市大学计算机学院) ; ADAPT Centre(ADAPT中心)
专题命中 视觉推理 :LLaVA(abstract);分类 cs.CV
机构 * Kyoto University(京都大学) ; The Hong Kong University of Science and Technology(香港科技大学) ; Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学) ; The University of Tokyo(东京大学)
专题命中 视觉推理 :visual reasoning(abstract);分类 cs.AI
Comments 9pages, 3 figures
机构 * University of Oxford(牛津大学) ; Qilu University of Technology(Shandong Academy of Sciences)(齐鲁工业大学(山东科学院)) ; Hong Kong University of Science and Technology(香港科学大学) ; Hong Kong University of Science and Technology (Guangzhou)(香港科学大学(广州))
专题命中 视觉推理 :InternVL(abstract);分类 cs.CV
机构 * University of California, Santa Cruz(加州大学圣克ruz分校) ; eBay(eBay公司)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI
机构 * DEVCOM Army Research Laboratory(DEVCOM陆军研究实验室) ; Lambda Inc(Lambda公司)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
Comments Project page: https://spatial-reasoner.github.io
机构 * University of Washington(华盛顿大学) ; Allen Institute for Artificial Intelligence(人工智能艾伦研究所) ; Stanford University(斯坦福大学) ; Woven by Toyota(丰田编织)
专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV
Comments CVPR 2025
机构 * KTH Royal Institute of Technology(皇家理工学院)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI
Comments 16 pages, 4 figures
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
Comments Dataset: https://huggingface.co/datasets/ivc-lrp/STSBench, Code: https://github.com/LRP-IVC/STSBench
机构 * CUHK MMLab(香港大学MMLab)
专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV
Comments Code is released at https://github.com/xinyan-cxy/MINT-CoT
机构 * University of Washington(华盛顿大学) ; Sun Yat-sen University(中山大学) ; Stanford University(斯坦福大学)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
Comments STARE is available at https://github.com/STARE-bench/STARE
机构 * ShanghaiTech University(上海科技大学) ; Microsoft Corporation(微软公司) ; Fudan University(复旦大学)
专题命中 视觉推理 :vision language model(abstract);分类 cs.CV
机构 * Carnegie Mellon University(卡内基梅隆大学) ; M-A-P ; Nanyang Technological University(南洋理工大学) ; University of Waterloo(滑铁卢大学) ; The University of Manchester(曼彻斯特大学)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
Comments ACL 2025 Main
机构 * National Tsing Hua University(清华大学)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV