Look Clearly Before Answering: Mitigating Hallucinations in LVLMs via Saliency-Driven Perceptual Realignment
回答前看清楚:通过显著性驱动的感知重新对齐减轻LVLMs中的幻觉
Pengxu Chen, Yao Zhu, Guangming Zhu, Jun Sheng, Jincai Huang, Xiangyang Ji, Liang Zhang
机构
*
Xidian University(西安电子科技大学)
;
Tsinghua University(清华大学)
;
Shanghai Road Transport Development Center(上海市道路运输发展中心)
;
Hunan Institute of Advanced Technology(湖南先进技术研究院)
机构
*
Tianjin Key Lab of Intelligent Unmanned Swarm Tech & System(天津智能集群技术与系统重点实验室)
;
Tianjin University(天津大学)
;
Institute of Computing and Intelligence(计算与智能研究所)
;
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
;
Tianjin Artificial Intelligence Innovation Center(天津人工智能创新中心)
;
Defense Innovation Institute Academy of Military Sciences(国防科技创新研究院)
;
School of Future Technology(未来技术学院)
;
Shanghai University(上海大学)
XRF-to-Optical Field-of-View Localization with Vision Language Models
基于视觉语言模型的X射线荧光(XRF)与光学显微镜视场(FOV)定位
Xiangyu Yin, Tatjana Paunesku, Letonia Copeland-Hardin, Martina Ralle, Zichao Wendy Di, Si Chen, Gayle E. Woloschak, Barry Lai, Mathew J. Cherukara, Stefan Vogt
机构
*
Northwestern University(西北大学)
;
University of Chicago(芝加哥大学)
;
Oregon Health and Science University(俄勒冈健康与科学大学)
;
Argonne National Laboratory(阿贡国家实验室)
专题命中
视觉定位与Grounding
:VLM(summary_cn,abstract);vision language model(title,abstract);分类 cs.CV
SPADE: Self-Play in Adaptive Synthetic Executable Environments
SPADE:自适应合成可执行环境中的自博弈
Bo Liu, Simon Yu, Yiding Jiang, Ao Qu, Andrew Zhao, Zichen Liu, Junsu Kim, Zijian Zhou, Seungone Kim, Tongzheng Ren, Mickel Liu, Hanfei Yu, Zhaorun Chen, Weiyan Shi, Paul Pu Liang, Luke Zettlemoyer, Yejin Choi, Natasha Jaques
机构
*
University of Washington(华盛顿大学)
;
Stanford University(斯坦福大学)
;
Northeastern University(东北大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Massachusetts Institute of Technology(麻省理工学院)
;
National University of Singapore(新加坡国立大学)
;
Seoul National University(首尔大学)
;
Stevens Institute of Technology(史蒂文斯理工学院)
;
University of Chicago(芝加哥大学)
COSTA: A Cluster-Centric Paradigm for Annotation-Free Open-Set Semantic Segmentation of Aerial Point Clouds with Domain Shifts
COSTA:面向存在域偏移的航拍点云无标注开放集语义分割的以聚类为中心范式
Yanghong Lin, Li Fang, Tianyu Li, Shudong Zhou, Wei Yao
机构
*
Fujian Institute of Research on the Structure of Matter, Chinese Academy of Sciences(中国科学院福建物质结构研究所)
;
Institute of Urban Environment, Chinese Academy of Sciences(中国科学院城市环境研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
School of Resource and Environmental Sciences, Wuhan University(武汉大学资源与环境科学学院)
Million-scale multimodal pollen microscopy with expert-guided foundation models
百万级多模态花粉显微镜图像与专家引导的基础模型
András Biricz, Björn Gedda, Donát Magyar, Antonio Spanu, János Fillinger, Péter Pollner, István Csabai
机构
*
Department of Physics of Complex Systems, ELTE Eötvös Loránd University(ELTE罗兰大学复杂物理系)
;
The Palynological Laboratory at the Swedish Museum of Natural History(瑞典自然历史博物馆孢粉学实验室)
;
National Centre for Public Health and Pharmacy(国家公共卫生与药品中心)
;
INRAE, UR 546 BioSP, Site Agroparc(法国国家农业、食品与环境研究院,UR 546 BioSP,阿格罗帕克园区)
;
National Korányi Institute for Pulmonology(国家科拉尼肺病研究所)
;
Health Data Science and AI Knowledge Centre, Health Services Management Training Centre, Faculty of Health and Public Administration, Semmelweis University(塞梅维什大学健康与公共管理学院卫生服务管理培训中心健康数据科学与人工智能知识中心)
;
Department of Biological Physics, ELTE Eötvös Loránd University(ELTE罗兰大学生物物理系)
AI总结
提出百万级多模态花粉显微镜数据集Pollen AI Atlas,结合专家引导的视觉-语言模型生成形态描述,实现跨区域、跨设置的高精度花粉识别与检索。
Comments31 pages, 5 main figures, supplementary information included. Submitted to Scientific Reports. v2: clarified reporting of taxonomic scope, captioning settings, backbone configuration, and evaluation details; no changes to numerical results or conclusions