HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models
Sushant Gautam, Michael A. Riegler, Pål Halvorsen
机构
*
Simula Metropolitan Center for Digital Engineering (SimulaMet)(Simula数字工程研究中心)
;
Oslo Metropolitan University (OsloMet)(奥斯陆 Metropolitan 大学)
;
Simula Research Laboratory(Simula研究实验室)
CoTBox-TTT: Grounding Medical VQA with Visual Chain-of-Thought Boxes During Test-time Training
Jiahe Qian, Yuhao Shen, Zhangtianyi Chen, Juexiao Zhou, Peisong Wang
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Data Science, The Chinese University of Hong Kong, Shenzhen(香港大学深圳研究院数据科学学院)
Tracing and Mitigating Hallucinations in Multimodal LLMs via Dynamic Attention Localization
Tiancheng Yang, Lin Zhang, Jiaye Lin, Guimin Hu, Di Wang, Lijie Hu
机构
*
MBZUAI
;
Provable Responsible AI and Data Analytics (PRADA) Lab(可证明负责任的人工智能与数据分析实验室)
;
King Abdullah University of Science and Technology(卡迪夫大学科学与技术大学)
;
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉科学学院)
;
University of Copenhagen(哥本哈根大学)
;
Tsinghua University(清华大学)
专题命中
视觉问答
:visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV
机构
*
Arizona State University(亚利桑那州立大学)
;
Purdue University(普渡大学)
;
University of British Columbia(不列颠哥伦比亚大学)
;
Stanford University(斯坦福大学)
;
The Ohio State University(俄亥俄州立大学)
;
Massachusetts General Hospital(麻省总医院)
;
Harvard Medical School(哈佛医学院)
;
Virginia Polytechnic Institute(弗吉尼亚多项技术学院)
;
University of Michigan(密歇根大学)
;
University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
;
Nanjing University(南京大学)
SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding
Zhenyu Yang, Yuhang Hu, Zemin Du, Dizhan Xue, Shengsheng Qian, Jiahong Wu, Fan Yang, Weiming Dong, Changsheng Xu
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Kuaishou Technology(快手科技)
;
Zhengzhou University(郑州大学)
;
ShanghaiTech University(上海科技大学)
;
Peng Cheng Laboratory(鹏城实验室)
Moving Pictures of Thought: Extracting Visual Knowledge in Charles S. Peirce's Manuscripts with Vision-Language Models
Carlo Teo Pedretti, Davide Picca, Dario Rodighiero
机构
*
Department of Classics, University Sapienza of Rome(罗马大学萨皮恩扎文学院)
;
Department of Language and Communication Sciences, University of Lausanne(洛桑大学语言与沟通科学系)
;
Campus Fryslân, University of Groningen(格罗宁根大学弗里斯兰校区)
专题命中
视觉推理
:vision-language model(title);VLM(abstract);visual language model(abstract);分类 cs.AI、cs.LG
Adaptive Diagnostic Reasoning Framework for Pathology with Multimodal Large Language Models
Yunqi Hong, Johnson Kao, Liam Edwards, Nein-Tzu Liu, Chung-Yen Huang, Alex Oliveira-Kowaleski, Cho-Jui Hsieh, Neil Y. C. Lin
机构
*
Computer Science Department, University of California, Los Angeles, CA, USA(加州大学洛杉矶分校计算机科学系)
;
Mechanical and Aerospace Engineering Department, University of California, Los Angeles, CA, USA(加州大学洛杉矶分校机械与航空航天工程系)
;
Department of Pathology, Tri-Service General Hospital, National Defense Medical Center, Taipei, Taiwan(台湾国防医学院三军总医院病理部)
;
Department of Pathology, National Taiwan University Hospital, Taipei, Taiwan(台湾国立台湾大学医院病理部)
;
Department of Pathology, David Geffen School of Medicine, University of California, Los Angeles, CA, USA(加州大学洛杉矶分校大卫·Geffen医学院病理部)
;
Bioengineering Department, University of California, Los Angeles, CA, USA(加州大学洛杉矶分校生物工程系)
;
Institute for Quantitative and Computational Biosciences, University of California, CA, USA(加州大学定量与计算生物科学研究所)
专题命中
视觉推理
:multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
TopoPerception: A Shortcut-Free Evaluation of Global Visual Perception in Large Vision-Language Models
Wenhao Zhou, Hao Zheng, Rong Zhao
机构
*
Center for Brain-Inspired Computing Research (CBICR)(脑启发计算研究中心)
;
Department of Precision Instruments(精密仪器系)
;
IDG/McGovern Institute for Brain Research(IDG/麦戈文脑研究学院)
GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory
Jeong Hun Yeo, Sangyun Chung, Sungjune Park, Dae Hoe Kim, Jinyoung Moon, Yong Man Ro
机构
*
Integrated Vision and Language Lab., School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST)(整合视觉与语言实验室,电气工程学院,韩国科学技术院(KAIST))
;
Visual Intelligence Research Section, Superintelligence Creative Research Laboratory, Electronics and Telecommunications Research Institute (ETRI)(视觉智能研究部,超智能创意研究实验室,电子电信研究院)
专题命中
视觉推理
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
ViSS-R1: Self-Supervised Reinforcement Video Reasoning
Bo Fang, Yuxin Song, Qiangqiang Wu, Haoyuan Sun, Wenhao Wu, Antoni B. Chan
机构
*
City University of Hong Kong(香港城市大学)
;
Baidu Inc.(百度公司)
;
Tsinghua University(清华大学)
;
The University of Sydney(悉尼大学)
专题命中
视觉推理
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
CommentsOur paper was initially titled "Video-SSR1: Self-Supervised Reinforcement Video Reasoning." Upon noticing its close resemblance to the title of a recently released paper, we have decided to rename our work as "ViSS-R1."