arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-29 至 2025-10-29 共收录 4 信号源:cs.CV, cs.AI, cs.LG

1. 幻觉与鲁棒性 4 篇

2510.17191 2025-10-29 cs.RO cs.AI 83%

SimpleVSF: VLM-Scoring Fusion for Trajectory Prediction of End-to-End Autonomous Driving

Peiru Zheng, Yun Zhao, Zhan Gong, Hong Zhu, Shaohua Wu

机构 * IEIT Systems(IEIT系统)

专题命中 幻觉与鲁棒性 :VLM(title,abstract);vision-language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24446 2025-10-29 cs.CL cs.CV 57%

SPARTA: Evaluating Reasoning Segmentation Robustness through Black-Box Adversarial Paraphrasing in Text Autoencoder Latent Space

Viktoriia Zinkovich, Anton Antonov, Andrei Spiridonov, Denis Shepelev, Andrey Moskalenko, Daria Pugacheva, Elena Tutubalina, Andrey Kuznetsov, Vlad Shakhuro

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11842 2025-10-29 cs.CV cs.CL 57%

Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs

Xuannan Liu, Zekun Li, Zheqi He, Peipei Li, Shuhan Xia, Xing Cui, Huaibo Huang, Xi Yang, Ran He

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) University of California, Santa Barbara(加州大学圣芭芭拉分校) Center for Research on Intelligent Perception and Computing, NLPR, CASIA(中国科学院CASIA智能感知与计算中心)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025 Dataset and Benchmark Track, Project page: https://liuxuannan.github.io/Video-SafetyBench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18672 2025-10-29 cs.CV 57%

CalFuse: Multi-Modal Continual Learning via Feature Calibration and Parameter Fusion

Juncen Guo, Siao Liu, Xiaoguang Zhu, Lianlong Sun, Liangyu Teng, Jingyi Wu, Di Li, Linxiao Gong, Weiwei Jiang, Wei Zhou, Liang Song

机构 * College of Intelligent Robotics and Advanced Manufacturing, Fudan University(智能机器人与先进制造学院,复旦大学) School of Future Science and Engineering, Soochow University(未来科学与工程学院,苏州大学) DataLab: Data Science and Informatics, University of California, Davis(数据实验室:数据科学与信息学,加州大学戴维斯分校) University of Rochester(罗切斯特大学) Ningbo University(宁波大学) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Beijing University of Posts and Telecommunications(北京邮电大学) Academy for Computer Science and Informatics, Cardiff University(计算机科学与信息学学院,卡迪夫大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏