Object-X: Learning to Reconstruct Multi-Modal 3D Object Representations
机构 * ETH Zurich(苏黎世联邦理工学院) ; Google(谷歌) ; Microsoft(微软)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * ETH Zurich(苏黎世联邦理工学院) ; Google(谷歌) ; Microsoft(微软)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) ; University of Oxford(牛津大学) ; Beijing Academy of Artificial Intelligence(北京人工智能研究院) ; The Hong Kong University of Science and Technology(香港科学与技术大学) ; Beijing University of Technology(北京工业大学) ; Tsinghua University(清华大学) ; City University of Hong Kong(香港城市大学)
专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV
Comments This paper is accepted by IJCAI2025 Workshop on Deepfake Detection, Localization, and Interpretability as Best Student Paper
机构 * Adobe Inc.(Adobe公司)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
Comments Accepted by 6th Workshop on Data Science with Human in the Loop @ VLDB 2025
机构 * Department of Systems Innovation, Graduate School of Engineering Science, Osaka University(大阪大学系统创新部门,工学研究科) ; OMRON SINIC X Corporation(OMRON SINIC X公司) ; The National Institute of Advanced Industrial Science and Technology (AIST)(国家先进工业科学与技术研究院)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG
机构 * Department of Electrical Engineering Stanford University(电气工程系 斯坦福大学) ; Emissary Technologies
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG
Comments Under review at ICLR 2026
机构 * Carnegie Mellon University(卡内基梅隆大学)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
Comments ICCV 2025 Accepted Paper
机构 * School of Computer Science and Engineering, University of Electronic Science and Technology of China(计算机科学与工程学院,电子科学与技术大学) ; Southwestern University of Finance and Economics(西南财经大学) ; Tongji University(同济大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG
机构 * The University of Tokyo(东京大学)
专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV
Comments ICCV2025 Workshop
机构 * Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) ; Artificial Intelligence Academy, Xidian University(西安电子科技大学人工智能学院)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
Comments 23 pages, 5 figures
机构 * NUS(新加坡国立大学) ; HKUST(GZ)(香港科技大学(广州)) ; NTU(南洋理工大学) ; HKUST(香港科技大学) ; I 2 R, A*STAR(I2R, A*STAR) ; IPAL, CNRS IRL 2955, Singapore(IPAL, CNRS IRL 2955, 新加坡) ; CerCo, CNRS UMR 5549, Université Toulouse III(CerCo, CNRS UMR 5549, 法国图卢兹第三大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV
Comments NeurIPS 2025 Spotlight; 43 pages, 17 figures, 16 tables; Project Page at https://talk2event.github.io
机构 * University of Nebraska Omaha(内布拉斯加大学奥马哈分校)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
Comments This version corrects the review of tau for negated atoms, and clarifies the distinction between global and local variables in conditional literals (the supporting proofs are also updated accordingly)
Journal ref In Practical Aspects of Declarative Languages: 27th International Symposium, PADL 2025, Denver, CO, USA, January 20-21, 2025, Proceedings. Springer-Verlag, Berlin, Heidelberg, 71-87
机构 * School of Informatics, Xiamen University(厦门大学信息学院)
专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV
Comments Accepted by ACM MM 2025. Project Page: https://jiajinglin.github.io/Phys4DGen
机构 * Thomson Reuters Foundational Research(汤姆森·路透基础研究) ; Thomson Reuters Labs(汤姆森·路透实验室) ; Imperial College London(帝国理工学院伦敦分校)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
Comments Accepted by 2025 EMNLP industry track
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
机构 * Department of Computer Science, University of California, Irvine(加州大学尔湾分校计算机科学系) ; School of Medicine, University of California, Irvine(加州大学尔湾分校医学院)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
机构 * Dept. of Electrical and Computer Engineering at Texas A&M University(德克萨斯A&M大学电气与计算机工程系) ; GRASP Laboratory at the University of Pennsylvania(宾夕法尼亚大学GRASP实验室)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
机构 * AMAP, Alibaba Group(阿里集团AMAP) ; Beijing University of Posts and Telecommunications(北京邮电大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV
机构 * Key Laboratory of Multimedia Trusted Perception(多媒体可信感知关键实验室) ; Efficient Computing, Ministry of Education of China, Xiamen University(高效计算、教育部中国 ministry of education、厦门大学) ; Institute of Artificial Intelligence, Xiamen University(人工智能研究院、厦门大学) ; Peng Cheng Laboratory, Shenzhen, China(鹏城实验室、深圳中国)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG
机构 * Tencent Youtu Lab(腾讯优图实验室)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
Comments 12 pages, 7 figures
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
Comments EMNLP 2025 Oral; Project Homepage: https://yanzehong.github.io/trust-vl/
机构 * Korea University(韩国大学) ; Korea Advanced Institute of Science and Technology(韩国科学技术院)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
Comments Accepted by NeurIPS 2025
机构 * Computational Bioscience Program University of Colorado Anschutz Medical Campus(科学生物学程序,科罗拉多大学安舒茨医学校区) ; Department of Pediatrics University of Chicago(儿科学系,芝加哥大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV
Comments Preprint under review at IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2025
机构 * Xiucheng Zhang(未知) ; Yang Jiang(未知) ; Jiashuo Bai(未知)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG
Comments 8 pages
机构 * NVIDIA ; Stanford University(斯坦福大学)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.AI
机构 * Institute of New Media and Communications(新媒体与通讯研究所) ; Dept. of Electrical and Computer Engineering(电气与计算机工程系)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
Comments Accepted to FM4RoboPlan workshop at RSS 2025
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
机构 * Zhejiang University(浙江大学) ; School of Engineering, Westlake University(西湖大学工程学院) ; Peking University(北京大学) ; Institute of Advanced Technology, Westlake Institute for Advanced Study(西湖先进研究院技术研究所)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
Comments Preprint version
机构 * School of Computer Science, University of South China(南方大学计算机科学学院) ; New Laboratory of Pattern Recognition, MAIS, CASIA(模式识别新实验室,MAIS,CASIA) ; School of Intelligence Science and Technology, Nanjing University(智能科学与技术学院,南京大学) ; Department of Computer Science and Engineering, University of California, Merced(加州大学默塞德分校计算机科学与工程系) ; Department of Computer Science and Engineering, Yonsei University(延世大学计算机科学与工程系)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
Comments Accepted by TPAMI