X-Reflect: Cross-Reflection Prompting for Multimodal Recommendation
机构 * University of Rochester(罗切斯特大学) ; Adobe Research(Adobe研究)
专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * University of Rochester(罗切斯特大学) ; Adobe Research(Adobe研究)
专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV
机构 * Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science & Technology, Beijing Institute of Technology, China(北京智能信息科技重点实验室,计算机科学与技术学院,北京理工大学,中国) ; Guangdong Provincial Key Laboratory of Machine Perception and Intelligent Computing, Shenzhen MSU-BIT University, Shenzhen, China(广东省机器感知与智能计算重点实验室,深圳MSU-BIT大学,深圳,中国) ; NVIDIA
专题命中 其他VLM :vision language model(abstract);分类 cs.CV
Comments NeurIPS 2025, Project: https://github.com/YaoChengTang/3D-Visual-Illusion-Depth-Estimation
机构 * Shanghai Jiao Tong University(上海交通大学) ; Nanyang Technological University(南洋理工大学)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
Comments Accepted at MM 2025
专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV
专题命中 其他VLM :vision-language model(abstract);分类 cs.LG
机构 * Meta ; University of Southern California(南加州大学)
专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV
Comments Accepted by NeurIPS 2025
机构 * School of Electrical and Computer Engineering, The University of Sydney(电气与计算机工程学院,悉尼大学) ; School of Computer Science, The University of Adelaide(计算机科学学院,阿德莱德大学) ; School of Computing and Information Technology, University of Wollongong(计算与信息科技学院,沃伦冈大学)
专题命中 其他VLM :vision language model(abstract);分类 cs.CV
机构 * Department of Computer Science, Hong Kong Baptist University(香港 Baptist 大学计算机科学系) ; NVIDIA AI Technology Center, NVIDIA(NVIDIA 人工智能技术中心)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
Comments Accepted at NeurIPS 2025
机构 * LLM Department, Tencent(腾讯大语言模型部门) ; Peking University(北京大学) ; Nanjing University(南京大学)
专题命中 其他VLM :MLLM(abstract);分类 cs.LG
机构 * University of Pennsylvania(宾夕法尼亚大学) ; Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
机构 * CSSE Department, Auburn University(计算机科学与工程系,阿伯丁大学)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
Comments Accepted to ICCV 2025. 17 pages, 6 figures, 3 tables
机构 * Max Planck Institute for Informatics(马克斯·普朗克研究所信息学研究所) ; EPFL(瑞士联邦理工学院)
专题命中 其他VLM :vision language model(abstract);分类 cs.LG
Comments SIGGRAPH Asia 2024 Courses. arXiv admin note: text overlap with arXiv:2208.11970 by other authors
Journal ref SIGGRAPH Asia 2024 Courses, Article No.: 8, Pages 1 - 27
专题命中 其他VLM :vision-language model(abstract);分类 cs.AI
机构 * Computing Department of Hong Kong Polytechnic University(香港理工大学计算机系) ; Computing Department, The Hong Kong Polytechnic University(香港理工大学计算机系)
专题命中 其他VLM :vision-language model(abstract);分类 cs.AI
Comments 13 pages, 8 figures. Submitted to IEEE Transactions on Information Forensics & Security
机构 * University of Science and Technology of China(中国科学技术大学) ; Suzhou Institute of Biomedical Engineering and Technology(苏州生物医学工程与技术研究所) ; Chinese Academy of Sciences(中国科学院) ; The Third Affiliated Hospital of Sun Yat-sen University(中山大学第三附属医院)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
Comments ICLR2026 under review
机构 * CyberAgent Tokyo Japan(CyberAgent东京日本)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
Comments This is the author's version of the work. It is posted here for your personal use. Not for redistribution. The definitive Version of Record was published in Proceedings of the 33rd ACM International Conference on Multimedia (MM '25), October 27-31, 2025, Dublin, Ireland, https://doi.org/10.1145/3746027.3758297
机构 * Indian Institute of Technology Patna(印度理工学院帕纳布分校) ; Sardar Patel Institute of Technology(萨达尔·帕特尔技术学院) ; Universitas Gadjah Mada(加查马大学) ; King Mongkut’s Institute of Technology Ladkrabang(拉差班国王技术学院) ; Shenzhen Technology University(深圳技术大学) ; Université de Toulouse(图卢兹大学)
专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI
Comments 52 pages, 56 figures; appearing at EMNLP'25
机构 * Qualcomm AI Research(高通人工智能研究)
专题命中 其他VLM :vision-language model(abstract);分类 cs.LG
机构 * The Chinese University of Hong Kong(香港中文大学) ; Columbia University in the City of New York(哥伦比亚大学) ; Singapore Management University(新加坡管理学院)
专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI
机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(基础模型研究中心、自动化研究所、中国科学院) ; Peng Cheng Laboratory, Shenzhen, China(鹏城实验室、深圳中国) ; School of Artificial Intelligence, University of Chinese Academy of Science, Beijing, China(人工智能学院、中国科学院大学、北京中国) ; Wuhan AI Research, Wuhan, China(武汉人工智能研究、武汉中国) ; MAPLE Lab, Westlake University(MAPLE实验室、西湖大学)
专题命中 其他VLM :vision language model(abstract);分类 cs.CV
机构 * Stability AI ; Arizona State University(亚利桑那州立大学) ; Google DeepMind(谷歌DeepMind)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
Comments NeurIPS 2025. Project Page : https://stable-cinemetrics.github.io/
机构 * The College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(南京航空航天大学人工智能学院) ; The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education(教育部脑机智能技术重点实验室)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
机构 * Teamreboott Inc.(Teamreboott公司) ; MIRI D.I.H Inc.(MIRI D.I.H公司)
专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV
机构 * Huazhong University of Science and Technology(华中科技大学)
专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV
Comments 4 figures, 4 tables
机构 * Key Laboratory of Pervasive Computing, Tsinghua University(清华大学普适计算重点实验室) ; The University of New South Wales(新南威尔士大学) ; SKLSDE Lab, Beihang University(北航SKLSDE实验室)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) ; University of Central Florida(中央佛罗里达大学) ; Cisco Research(思科研究)
专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV
专题命中 其他VLM :MLLM(abstract);分类 cs.AI
Comments 31 pages, 16 figures, 12 tables
机构 * Department of Pathology and Laboratory Medicine, University of Pennsylvania(病理学与实验室医学系,宾夕法尼亚大学) ; Department of Electrical and System Engineering, University of Pennsylvania(电气与系统工程系,宾夕法尼亚大学) ; The Wharton School, University of Pennsylvania(沃顿商学院,宾夕法尼亚大学) ; Department of Bioengineering, University of Pennsylvania(生物工程系,宾夕法尼亚大学) ; Department of Computer and Information Science, University of Pennsylvania(计算机与信息科学系,宾夕法尼亚大学) ; Department of Biostatistics, Epidemiology & Informatics, University of Pennsylvania(生物统计学、流行病学与信息学系,宾夕法尼亚大学)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
机构 * Korea University(韩国大学) ; Hanwha Vision(翰威英航) ; KAIST(韩国科学技术院)
专题命中 其他VLM :MLLM(abstract);分类 cs.CV
Comments EMNLP 2025 Findings
机构 * Australian Institute for Machine Learning, University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学) ; Stony Brook University(石溪大学)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV