arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-01-15 至 2026-01-15 共收录 3 信号源:cs.CV, cs.AI, cs.LG

1. 幻觉与鲁棒性 3 篇

2601.02316 2026-01-15 cs.LG cs.AI 84%

DatBench: Discriminative, Faithful, and Efficient VLM Evaluations

DatBench:具有辨别性、忠实性和效率的视觉语言模型评估

DatologyAI, :, Siddharth Joshi, Haoli Yin, Rishabh Adiga, Ricardo Monti, Aldo Carranza, Alex Fang, Alvin Deng, Amro Abbas, Brett Larsen, Cody Blakeney, Darren Teh, David Schwab, Fan Pan, Haakon Mongstad, Jack Urbanek, Jason Lee, Jason Telanoff, Josh Wills, Kaleigh Mentzer, Luke Merrick, Parth Doshi, Paul Burstein, Pratyush Maini, Scott Loftin, Spandan Das, Tony Jiang, Vineeth Dorna, Zhengping Wang, Bogdan Gaza, Ari Morcos, Matthew Leavitt

专题命中 幻觉与鲁棒性 :VLM(title,abstract);vision-language model(abstract);分类 cs.AI、cs.LG

AI总结 DatBench通过转换和过滤现有基准,提升视觉语言模型评估的忠实性和效率,实现13倍加速并保持辨别性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08860 2026-01-15 cs.CV cs.AI 81%

Bias Detection and Rotation-Robustness Mitigation in Vision-Language Models and Generative Image Models

视觉-语言模型和生成图像模型中的偏见检测与旋转鲁棒性缓解

Tarannum Mithila

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出旋转鲁棒缓解策略,通过数据增强、表征对齐和模型正则化,提升视觉-语言和生成图像模型在旋转和分布偏移下的鲁棒性和公平性。

Comments Preprint. This work is derived from the author's Master's research. Code and supplementary materials will be released separately

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09165 2026-01-15 cs.LG 57%

Multi-Teacher Ensemble Distillation: A Mathematical Framework for Probability-Domain Knowledge Aggregation

多教师集成蒸馏:概率域知识聚合的数学框架

Aaron R. Flouro, Shawn P. Chadwick

机构 * Sparse-Tech

专题命中 幻觉与鲁棒性 :grounding(abstract);分类 cs.LG

AI总结 本文提出了一种概率域知识聚合的数学框架,通过公理化方法定义了多教师集成蒸馏的核心原理,并提供了理论保障和多种实现策略。

Comments 7 pages, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏