I2I-STRADA -- Information to Insights via Structured Reasoning Agent for Data Analysis
机构 * Mphasis Limited(默比斯有限公司)
专题命中 视觉推理 :grounding(abstract);分类 cs.AI
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * Mphasis Limited(默比斯有限公司)
专题命中 视觉推理 :grounding(abstract);分类 cs.AI
机构 * The University of Hong Kong(香港大学) ; Huawei Noah’s Ark Lab(华为诺亚实验室)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
Comments 37 pages, 12 figures
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) ; Microsoft(微软) ; UIUC(伊利诺伊大学香槟分校)
专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV
机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) ; Nanjing University of Science and Technology(南京理工大学)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
Comments Code: https://dongzhang89.github.io/RGenie.github.io/
机构 * Tsinghua University(清华大学) ; Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ; ByteDance China(字节跳动中国) ; Peng Cheng Laboratory(鹏城实验室) ; Wuhan AI Research(武汉人工智能研究)
专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV
机构 * Department of Computer Science, Norwegian University of Science and Technology(计算机科学系,挪威科学技术大学) ; Department of Design, Norwegian University of Science and Technology(设计系,挪威科学技术大学)
专题命中 视觉推理 :visual reasoning(abstract);分类 cs.AI
机构 * Halıcıoğlu Data Science Institute (HDSI), University of California San Diego, La Jolla, CA, USA(Halıcıoğlu数据科学研究所(HDSI),加州大学圣地亚哥分校,拉贾拉,加州,美国)
专题命中 视觉推理 :VLM(abstract);分类 cs.AI
Comments 8 pages, ICML MAS workshop
机构 * Shanxi Province Cancer Hospital, Chinese Academy of Medical Sciences(山西省肿瘤医院、中国医学科学院) ; Suzhou Institute of Biomedical Engineering and Technology, Chinese Academy of Sciences(苏州生物医学工程与技术研究所、中国科学院)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.LG
机构 * Sun Yat-sen University, Shenzhen(中山大学深圳校区) ; Chinese University of Hong Kong, Shenzhen(香港中文大学深圳校区) ; SenseTime Research(商汤科技研究院) ; Chinese University of Hong Kong(香港中文大学) ; Guangdong Key Laboratory of Big Data Analysis and Processing(广东省大数据分析与处理重点实验室)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
Comments Extended Version of KptLLM. arXiv admin note: text overlap with arXiv:2411.01846
机构 * University of Science and Technology of China(中国科学技术大学) ; YuanShi Technology(元世科技) ; Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身人工智能研究院)
专题命中 视觉推理 :MLLM(abstract);分类 cs.CV
Comments ICCV 2025
机构 * Massachusetts Institute of Technology(麻省理工学院) ; University of Alabama(阿拉巴马大学)
专题命中 视觉推理 :grounding(abstract);分类 cs.AI
Comments Accepted as paper in 19th International Conference on Neurosymbolic Learning and Reasoning,NeSy 2025
机构 * Baichuan Inc.(北京百川科技有限公司) ; Peking University(北京大学)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI
Comments 14 pages,4 figures
机构 * Center for research in Computer Vision, University of Central Florida(计算机视觉研究中心,中央佛罗里达大学) ; Cornell University(康奈尔大学) ; Department of Electronics Engineering, Mehran University of Engineering & Technology(电子工程系,工程与技术大学) ; Department of Imaging Physics, The University of Texas MD Anderson Cancer Center(成像物理系,德克萨斯大学MD安德森癌症中心) ; Intelligent Transportation Systems, University of Tennessee(智能交通运输系统,田纳西大学) ; Department of Electrical Engineering, City University of Hong Kong(电气工程系,香港城市大学) ; Department of Mechanical Engineering, Bartin University(机械工程系,巴廷大学) ; Vector Institute, Toronto Canada(多伦多加拿大向量研究所) ; Manchester Metropolitan University(曼彻斯特 Metropolitan 大学) ; Center for Data Science, New York University(数据科学中心,纽约大学) ; Meta Research(Meta 研究) ; Amazon Research(亚马逊研究) ; Centre for Artificial Intelligence Research and Optimization, Torrens University Australia(人工智能研究与优化中心,塔伦斯大学澳大利亚) ; University Research and Innovation Center, Obuda University(研究与创新中心,奥布达大学)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI
机构 * Institutes of Innovation for Future Society, Nagoya University, Japan(面向未来的创新研究所,名古屋大学,日本)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
机构 * MBZUAI ; The University of Sydney(悉尼大学) ; AI2Robotic ; Texas A&M University(德克萨斯大学) ; The University of Melbourne(墨尔本大学)
专题命中 视觉推理 :grounding(abstract);分类 cs.CV
机构 * The Laboratory of Intelligent Collaborative Computing of UESTC(UESTC智能协同计算实验室) ; Ubiquitous Intelligence and Trusted Services Key Laboratory of Sichuan Province(四川省 Ubiquitous Intelligence and Trusted Services 重点实验室) ; Monash University(墨尔本大学) ; Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东省人工智能与数字经济实验室(深圳))
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
Comments Accepted to ICCV 2025
机构 * Xiaoduo AI(小多AI) ; University of Dayton(代顿大学)
专题命中 视觉推理 :MLLM(abstract);分类 cs.AI
机构 * ECE, Seoul National University, Korea.(电子工程系,首尔国立大学,韩国) ; IPAI, Seoul National University, Korea.(人工智能研究所,首尔国立大学,韩国)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
Comments Accepted to the 42nd International Conference on Machine Learning (ICML 2025)
机构 * Fraunhofer Institute for Integrated Circuits IIS(弗劳恩霍夫集成电路研究所)
专题命中 视觉推理 :LLaVA(abstract);分类 cs.AI
Journal ref IEEE Wireless Communications and Networking Conference (WCNC), March 2025, Milan, Italy
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
Comments Technical Report: https://github.com/Kwai-Keye/Keye
专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI
机构 * School of Computing and Augmented Intelligence, Arizona State University(计算与增强智能学院,亚利桑那州立大学)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI
机构 * HKUST(香港理工大学) ; Dartmouth College(达特茅斯学院)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
Comments Project page: https://danielshkao.github.io/thinkfirst.html
机构 * Department of Aerospace Engineering and Coordinated Science Laboratory, University of Illinois at Urbana-Champaign(航空航天工程系和协调科学实验室,伊利诺伊大学厄巴纳-香槟分校) ; Department of Electrical Engineering and Coordinated Science Laboratory, University of Illinois at Urbana-Champaign(电气工程系和协调科学实验室,伊利诺伊大学厄巴纳-香槟分校) ; Intelligent Robotics Group, NASA Ames Research Center(智能机器人组,美国国家航空航天局阿姆斯研究中心)
专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV
Comments 4 pages, 5 figures, presented at the Workshop on 3D Visual Representations for Manipulation at the 2023 IEEE International Conference on Robotics and Automation in Yokohama, Japan. Video presentation [https://youtu.be/mg30uCUtpOk]. Poster [https://hollydinkel.github.io/assets/pdf/ICRA20243DVRM_poster.pdf] 3DVRM Workshop [https://3d-manipulation-workshop.github.io/]
机构 * Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
机构 * LMU Munich(慕尼黑大学) ; Munich Center for Machine Learning(慕尼黑机器学习中心) ; University of Oxford(牛津大学)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
Comments preprint version; 23 pages (including references and appendix)
机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(中国教育部多媒体可信感知与高效计算重点实验室,厦门大学) ; The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; The Chinese University of Hong Kong(香港中文大学) ; Bytedance(字节跳动) ; National University of Singapore(新加坡国立大学) ; Tsinghua University(清华大学)
专题命中 视觉推理 :MLLM(abstract);分类 cs.CV
Comments 40 pages, 26 figures
机构 * The Hong Kong Polytechnic University(香港理工大学) ; Zhejiang University(浙江大学) ; University of Electronic Science and Technology of China(电子科技大学) ; Reallm Labs(Reallm 实验室) ; Amazon(亚马逊) ; The Hong Kong University of Science and Technology(香港理工大学)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI