Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision
共收录 11875 篇
2603.190262026-03-20cs.CV
Rethinking MLLM Itself as a Segmenter with a Single Segmentation Token
重新思考MLLM本身作为分割器:仅用一个分割标记
Anqi Zhang, Xiaokang Ji, Guangyu Gao, Jianbo Jiao, Chi Harold Liu, Yunchao Wei
机构
*
Beijing Institute of Technology(北京理工大学)
;
University of Birmingham(伯明翰大学)
;
Beijing Jiaotong University(北京交通大学)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
机构
*
Intelligent Software Research Center, Institute of Software, CAS(软件研究所智能软件研究中心,中国科学院)
;
State Key Lab of Processors, Institute of Computing Technology, CAS(中国科学院计算技术研究所处理器国家重点实验室)
;
School of Information Science and Technology, University of Science and Technology of China(中国科学技术大学信息科学与技术学校)
All-in-One Slider for Attribute Manipulation in Diffusion Models
用于扩散模型属性操控的全能滑块
Weixin Ye, Hongguang Zhu, Wei Wang, Yahui Liu, Mengyu Wang, Xuecheng Nie
机构
*
Institute of Information Science, Beijing Jiaotong University(北京交通大学信息科学学院)
;
Visual Intelligence + X International Cooperation Joint Laboratory of the Ministry of Education(教育部视觉智能+X国际合作联合实验室)
;
City University of Macau(澳门城市大学)
;
Kuaishou(快手)
;
Meitu(美图)
Look Before You Fuse: 2D-Guided Cross-Modal Alignment for Robust 3D Detection
先看再融合:基于2D引导的跨模态对齐用于鲁棒的3D检测
Xiang Li, Zhangchi Hu, Xiao Xu, Bin Kong
机构
*
Institute of Intelligent Machines, Hefei Institutes of Physical Science, Chinese Academy of Sciences(智能机器研究所,合肥物理科学研究院,中国科学院)
;
University of Science and Technology of China(中国科学技术大学)
TINA: Text-Free Inversion Attack for Unlearned Text-to-Image Diffusion Models
TINA:无文本逆向攻击用于未学习的文本到图像扩散模型
Qianlong Xiang, Miao Zhang, Haoyu Zhang, Kun Wang, Junhui Hou, Liqiang Nie
机构
*
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
;
City University of Hong Kong(香港城市大学)
;
Shenzhen Loop Area Institute(深圳河套学院)
;
Peng Cheng Laboratory(鹏城实验室)
;
Shandong University(山东大学)
机构
*
State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(多模态人工智能系统国家重点实验室,中国科学院自动化所)
;
School of Artificial Intelligence, UCAS(人工智能学院,中国科学院大学)
;
Beijing National Research Center for Information Science and Technology(北京信息科学研究中心)
;
Institute of Artificial Intelligence, USTB(信息科学技术大学人工智能学院)
;
School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院)
;
Zhongguancun Academy(中关村学院)
Does YOLO Really Need to See Every Training Image in Every Epoch?
YOLO真的需要在每个epoch中都看到每张训练图像吗?
Xingxing Xie, Jiahua Dong, Junwei Han, Gong Cheng
机构
*
School of Automation, Northwestern Polytechnical University(西北工业大学自动化学院)
;
School of Artificial Intelligence, Chongqing University of Posts and Telecommunications(重庆邮电大学人工智能学院)
FINER: MLLMs Hallucinate under Fine-grained Negative Queries
FINER:MLLMs在细粒度负查询下产生幻觉
Rui Xiao, Sanghwan Kim, Yongqin Xian, Zeynep Akata, Stephan Alaniz
机构
*
Technical University of Munich(慕尼黑技术大学)
;
Munich Center for Machine Learning(慕尼黑机器学习中心)
;
Helmholtz Munich(海德堡-慕尼黑亥姆霍尔茨中心)
;
Google(谷歌)
;
LTCI, Télécom Paris, Institut Polytechnique de Paris(LTCI,巴黎电信学院,巴黎理工学院)
机构
*
German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI))
;
RPTU University Kaiserslautern-Landau(科布伦茨-兰道大学(RPTU))
;
University of Modena and Reggio Emilia(摩德纳和雷吉奥-艾米利亚大学)