A-SEA3L-QA: A Fully Automated Self-Evolving, Adversarial Workflow for Arabic Long-Context Question-Answer Generation
机构 * Humain
专题命中 视觉问答 :vision language model(abstract);分类 cs.AI
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * Humain
专题命中 视觉问答 :vision language model(abstract);分类 cs.AI
专题命中 视觉推理 :grounding(abstract);multimodal large language model(abstract);分类 cs.CV
机构 * Syracuse University Syracuse New York USA ; Arizona State University Tempe Arizona USA ; Scientific Systems Company, Inc. Woburn Massachusetts USA ; Department of Computer Science ; Engineering, Universidad Nacional del Sur (UNS) \& Institute for Computer Science ; Syracuse University ; Arizona State University ; Scientific Systems Company, Inc.
专题命中 视觉推理 :grounding(abstract);分类 cs.AI、cs.LG
专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV
Comments Work under review in NeurIPS 2025 with the title "Are we using Motion in Referring Segmentation? A Motion-Centric Evaluation"
专题命中 视觉定位与Grounding :MLLM(title);multimodal large language model(abstract);分类 cs.CV
机构 * College of Software Technology, Zhejiang University(浙江大学软件技术学院) ; College of Computer Science, Zhejiang University(浙江大学计算机科学学院)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV
机构 * Duke University(杜克大学)
专题命中 视觉定位与Grounding :vision language model(title,abstract);分类 cs.CV
Comments The paper has been accepted to the 2025 IEEE Conference on Virtual Reality and 3D User Interfaces (IEEE VR), and selected for publication in the 2025 IEEE Transactions on Visualization and Computer Graphics (TVCG) special issue
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV
Comments Accepted by ACMMM2025
机构 * Enuma, Inc.(Enuma公司) ; Korea University(韩国大学)
专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract)
机构 * Media Arts & Technology UC Santa Barbara(媒体艺术与技术大学圣芭芭拉分校)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
Comments to be published in IEEE VISAP 2025
机构 * University of Glasgow(格拉斯哥大学) ; Jadavpur University(贾瓦德普尔大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
Comments I want to revisit some of the experiments in this paper, specifically figure 5
机构 * School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen)(计算机科学与技术学院,哈尔滨工业大学(深圳))
专题命中 GUI与屏幕智能体 :VLM(title,abstract);vision-language model(abstract);分类 cs.CV
Comments Project Page: https://github.com/JiuTian-VL/Large-VLM-based-VLA-for-Robotic-Manipulation
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; The Hong Kong University of Science and Technology(香港科技大学) ; Tsinghua University(清华大学) ; Ant Group(蚂蚁集团) ; Alibaba Group(阿里巴巴集团)
专题命中 幻觉与鲁棒性 :visual question answering(abstract);multimodal large language model(abstract);分类 cs.AI、cs.LG
机构 * Manning College of Information \& Computer Sciences, University of Massachusetts Amherst, Amherst, U.S.
专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);LLaVA(abstract);分类 cs.LG
机构 * ZERON Shanghai, China(上海零点科技有限公司)
专题命中 VLM训练与架构 :vision language model(title,abstract);VLM(abstract);分类 cs.CV
Comments 2nd place in CVPR 2024 End-to-End Driving at Scale Challenge
机构 * Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,中央佛罗里达大学)
专题命中 VLM训练与架构 :LLaVA(abstract);分类 cs.CV、cs.LG
Comments The dataset and code used in this submission is available at: https://ucf-crcv.github.io/GAEA/
机构 * Keio University(庆应大学) ; NVIDIA(英伟达)
专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Comments Accepted to ICCV Workshop 2025
机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University, China(清华大学深圳国际研究生院,清华大学,中国)
专题命中 其他VLM :vision-language model(abstract);分类 cs.AI、cs.LG
Comments Accepted by the European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML-PKDD 2025)