Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
VLMs 是在看还是只是在说?揭示视觉重新检查的幻觉
Chufan Shi, Cheng Yang, Yaokang Wu, Linghao Jin, Bo Shui, Taylor Berg-Kirkpatrick, Xuezhe Ma
机构
*
University of Southern California(南加州大学)
;
University of California San Diego(加州大学圣地亚哥分校)
;
Carnegie Mellon University(卡内基梅隆大学)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构
*
SIAT, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Alibaba Group(阿里巴巴集团)
;
University of New South Wales(新南威尔士大学)
专题命中
视觉定位与Grounding
:VLM(summary_cn);vision language model(abstract);分类 cs.AI
机构
*
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
;
Zhongguancun Academy(中关村学院)
;
Huazhong University of Science and Technology(华中科技大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
South China University of Technology(华南理工大学)
;
University of Oxford(牛津大学)
;
Peking University(北京大学)
AI总结
本文提出了一种三阶段的代理编程准备方法,通过情境 grounding、协作规范和任务分解,提升 AI 编程的效率与质量,通过 hackathon 实验验证了准备阶段的重要性。
Comments5 pages. Accepted at VibeX 2026, the 1st International Workshop on Vibe Coding and Vibe Researching, co-located with EASE 2026, Glasgow, June 9-12 2026. Camera-ready version. Research artifact: https://doi.org/10.5281/zenodo.19868258
ST-$π$: Structured SpatioTemporal VLA for Robotic Manipulation
ST-$π$: 结构化时空VLA用于机器人操作
Chuanhao Ma, Hanyu Zhou, Shihan Peng, Yan Li, Tao Gu, Luxin Yan
机构
*
School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院)
;
School of Computing, National University of Singapore(新加坡国立大学计算机学院)
;
School of Computing, Macquarie University(麦考瑞大学计算机学院)
All Changes May Have Invariant Principles: Improving Ever-Shifting Harmful Meme Detection via Design Concept Reproduction
所有变化可能都有不变原则:通过设计概念再现改进永动有害迷因检测
Ziyou Jiang, Mingyang Li, Junjie Wang, Yuekai Huang, Jie Huang, Zhiyuan Chang, Zhaoyang Li, Qing Wang
机构
*
State Key Laboratory of Complex System Modeling and Simulation Technology(复杂系统建模与仿真技术国家重点实验室)
;
Science and Technology on Integrated Information System Laboratory Institute of Software Chinese Academy of Sciences(软件研究所信息集成系统技术研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
专题命中
视觉定位与Grounding
:MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV
Specializing Large Models for Oracle Bone Script Interpretation via Component-Grounded Multimodal Knowledge Augmentation
通过组件导向的多模态知识增强专门化大型模型进行甲骨文解读
Jianing Zhang, Runan Li, Honglin Pang, Ding Xia, Zhou Zhu, Qian Zhang, Chuntao Li, Xi Yang
机构
*
College of Software, Jilin University(吉林大学软件学院)
;
School of Artificial Intelligence, Jilin University(吉林大学人工智能学院)
;
Graduate School of Information Science and Technology, The University of Tokyo(东京大学信息科学与技术研究生院)
;
School of Archaeology, Jilin University(吉林大学考古学院)
;
Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, MoE, China(教育部知识驱动人机智能工程研究中心)