Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
视频中的奉承:视频大语言模型中奉承行为的基准测试与分析
Wenrui Zhou, Mohamed Hendy, Shu Yang, Qingsong Yang, Zikun Guo, Yuyu Luo, Lijie Hu, Di Wang
机构
*
Provable Responsible AI and Data Analytics (PRADA) Lab(可证负责任人工智能与数据 analytics 实验室)
;
King Abdullah University of Science and Technology(国王阿卜杜勒阿齐兹科学与技术大学)
;
HKUST(香港科技大学)
;
MBZUAI(穆罕默德·本·拉希德人工智能研究所)
;
University of Science and Technology of China(中国科学技术大学)
;
Kyungpook National University(庆尚国立大学)
SurgiSR4K: A High-Resolution Endoscopic Video Dataset for Robotic-Assisted Minimally Invasive Procedures
SurgiSR4K:一种用于机器人辅助微创手术的高分辨率内窥镜视频数据集
Fengyi Jiang, Xiaorui Zhang, Lingbo Jin, Ruixing Liang, Yuxin Chen, Adi Chola Venkatesh, Jason Culman, Tiantian Wu, Lirong Shao, Wenqing Sun, Cong Gao, Hallie McNamara, Jingpei Lu, Omid Mohareri
机构
*
Intuitive Surgical, Inc.(Intuitive Surgical公司)
;
Johns Hopkins Medicine Neurosurgery(约翰霍普金斯医学神经外科)
;
Johns Hopkins University Electrical and Computer Engineering(约翰霍普金斯大学电气与计算机工程)
;
University of British Columbia Electrical and Computer Engineering(不列颠哥伦比亚大学电气与计算机工程)
;
Wilford & Kate Bailey Small Animal Teaching Hospital(威尔福德与凯蒂·贝利小动物教学医院)
CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding
CGC:基于组成性 grounded 对比的细粒度多图像理解
Lihao Zheng, Zhenwei Shao, Yu Zhou, Yan Yang, Xintian Shen, Jiawei Chen, Hao Ma, Tao Wei
机构
*
School of Computer Science and Technology, Hangzhou Dianzi University(杭州电子科技大学计算机科学与技术学院)
;
School of Computer Science(计算机科学学院)
;
Technology, Hangzhou Dianzi University(技术,杭州电子科技大学)
专题命中
视觉定位与Grounding
:grounding(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI
AD-Copilot: A Vision-Language Assistant for Industrial Anomaly Detection via Visual In-context Comparison
AD-Copilot:通过视觉上下文比较实现工业异常检测的视觉-语言助手
Xi Jiang, Yue Guo, Jian Li, Yong Liu, Bin-Bin Gao, Hanqiu Deng, Jun Liu, Heng Zhao, Chengjie Wang, Feng Zheng
机构
*
Department of Computer Science and Engineering, Southern University of Science and Technology(计算机科学与工程系,南方科技大学)
;
Macau University of Science and Technology(澳门科技大学)
;
Tencent YouTu Lab(腾讯YouTu实验室)
;
Nanjing University(南京大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Centre for Frontier AI Research, A*STAR(前沿人工智能研究中心,A*STAR)
专题命中
视觉定位与Grounding
:MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI
Fall into a Pit, Gain in a Wit: Cognitive-Guided Harmful Meme Detection via Misjudgment Risk Pattern Retrieval
跌入陷阱,获得智慧:通过误判风险模式检索的认知引导有害迷因检测
Wenshuo Wang, Ziyou Jiang, Junjie Wang, Mingyang Li, Jie Huang, Yuekai Huang, Zhiyuan Chang, Feiyan Duan, Qing Wang
机构
*
State Key Laboratory of Complex System Modeling and Simulation Technology(复杂系统建模与仿真技术国家重点实验室)
;
Science and Technology on Integrated Information System Laboratory(集成信息系统技术研究所)
;
Institute of Software Chinese Academy of Sciences(中国科学院软件研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
专题命中
视觉定位与Grounding
:MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.AI、cs.LG
Watch Before You Answer: Learning from Visually Grounded Post-Training
在回答前观看:从视觉引导的后训练中学习
Yuxuan Zhang, EunJeong Hwang, Huaisong Zhang, Penghui Du, Yiming Jia, Dongfu Jiang, Xuan He, Shenhui Zhang, Ping Nie, Peter West, Kelsey R. Allen
机构
*
University of British Columbia(不列颠哥伦比亚大学)
;
Vector Institute(向量研究所)
;
Etude AI
;
Kolors Team, Kuaishou Technology(快手科技Kolors团队)
;
University of Toronto(多伦多大学)
;
University of Waterloo(滑铁卢大学)
;
University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
Symphony: A Cognitively-Inspired Multi-Agent System for Long-Video Understanding
Symphony:一种受认知启发的多智能体系统用于长视频理解
Haiyang Yan, Hongyun Zhou, Peng Xu, Xiaoxue Feng, Mengyi Liu
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Kuaishou Technology(快手科技)
;
School of Future Technology, University of Chinese Academy of Sciences(中国科学院大学未来技术学院)
机构
*
Shanghai Key Lab of Intelligent Information Processing, College of Computer Science and Artificial Intelligence, Fudan University(上海智能信息处理关键实验室,计算机科学与人工智能学院,复旦大学)
;
Artificial Intelligence, Fudan University(人工智能,复旦大学)
专题命中
视觉定位与Grounding
:grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
机构
*
ARTORG Center for Biomedical Engineering Research, University of Bern, Switzerland(ARTORG生物医学工程研究中心,伯尔尼大学,瑞士)
;
Shanghai Jiao Tong University, China(上海交通大学,中国)
;
Kaiko.AI, Switzerland(Kaiko.AI,瑞士)
;
Dept. of Radiation Oncology, Inselspital, Bern University Hospital(放射肿瘤科,因斯普尔茨医院,伯尔尼大学医院)
Taming a Retrieval Framework to Read Images in Humanlike Manner for Augmenting Generation of MLLMs
Suyang Xi, Chenxi Yang, Hong Ding, Yiqing Ni, Catherine C. Liu, Yunhao Liu, Chengqi Zhang
机构
*
Emory University(埃默里大学)
;
University of Electronic Science and Technology of China(电子科技大学)
;
University of Illinois Chicago(伊利诺伊大学香槟分校)
;
The Hong Kong Polytechnic University(香港理工大学)
专题命中
视觉定位与Grounding
:visual question answering(abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
What is the Visual Cognition Gap between Humans and Multimodal LLMs?
Xu Cao, Yifan Shen, Bolin Lai, Wenqian Ye, Yunsheng Ma, Joerg Heintz, Jintai Chen, Meihuan Huang, Jianguo Cao, Aidong Zhang, James M. Rehg
机构
*
Department of Computer Science, University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机科学系)
;
College of Computing, Georgia Institute of Technology(佐治亚理工学院计算机学院)
;
Department of Computer Science, University of Virginia(弗吉尼亚大学计算机科学系)
;
Digital Twin Lab, Purdue University(普渡大学数字孪生实验室)
;
HKUST (Guangzhou)(香港科技大学(广州))
;
Department of Rehabilitation Medicine, Shenzhen Children’s Hospital(深圳儿童医院康复医学系)
专题命中
视觉定位与Grounding
:vision language model(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
Erik Daxberger, Nina Wenzel, David Griffiths, Haiming Gang, Justin Lazarow, Gefen Kohavi, Kai Kang, Marcin Eichner, Yinfei Yang, Afshin Dehghan, Peter Grasch
机构
*
Apple(苹果公司)
专题命中
视觉定位与Grounding
:grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.LG