arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-04 至 2025-09-04 共收录 10 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 10 篇

2503.15867 2025-09-04 cs.CV cs.AI 86%

TruthLens: Visual Grounding for Universal DeepFake Reasoning

Rohit Kundu, Shan Jia, Vishal Mohanty, Athula Balachandran, Amit K. Roy-Chowdhury

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02807 2025-09-04 cs.CV 79%

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?

Mennatullah Siam

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Work under review in NeurIPS 2025 with the title "Are we using Motion in Referring Segmentation? A Motion-Centric Evaluation"

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19493 2025-09-04 cs.CR cs.CV 79%

Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents

Zhixin Lin, Jungang Li, Shidong Pan, Yibo Shi, Yue Yao, Dongliang Xu

专题命中 视觉定位与Grounding :MLLM(title);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01197 2025-09-04 cs.CV cs.RO 79%

A Coarse-to-Fine Approach to Multi-Modality 3D Occupancy Grounding

Zhan Shi, Song Wang, Junbo Chen, Jianke Zhu

机构 * College of Software Technology, Zhejiang University(浙江大学软件技术学院) College of Computer Science, Zhejiang University(浙江大学计算机科学学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12553 2025-09-04 cs.CV 79%

ViDDAR: Vision Language Model-Based Task-Detrimental Content Detection for Augmented Reality

Yanming Xiu, Tim Scargill, Maria Gorlatova

机构 * Duke University(杜克大学)

专题命中 视觉定位与Grounding :vision language model(title,abstract);分类 cs.CV

Comments The paper has been accepted to the 2025 IEEE Conference on Virtual Reality and 3D User Interfaces (IEEE VR), and selected for publication in the 2025 IEEE Transactions on Visualization and Computer Graphics (TVCG) special issue

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01275 2025-09-04 cs.CV 70%

Novel Category Discovery with X-Agent Attention for Open-Vocabulary Semantic Segmentation

Jiahao Li, Yang Lu, Yachao Zhang, Fangyong Wang, Yuan Xie, Yanyun Qu

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

Comments Accepted by ACMMM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02915 2025-09-04 cs.CL 67%

English Pronunciation Evaluation without Complex Joint Training: LoRA Fine-tuned Speech Multimodal LLM

Taekyung Ahn, Hosung Nam

机构 * Enuma, Inc.(Enuma公司) Korea University(韩国大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02924 2025-09-04 cs.MM cs.AI cs.HC 57%

Simulacra Naturae: Generative Ecosystem driven by Agent-Based Simulations and Brain Organoid Collective Intelligence

Nefeli Manoudaki, Mert Toka, Iason Paterakis, Diarmid Flatley

机构 * Media Arts & Technology UC Santa Barbara(媒体艺术与技术大学圣芭芭拉分校)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments to be published in IEEE VISAP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02837 2025-09-04 cs.IR cs.AI 57%

HF-RAG: Hierarchical Fusion-based RAG with Multiple Sources and Rankers

Payel Santra, Madhusudan Ghosh, Debasis Ganguly, Partha Basuchowdhuri, Sudip Kumar Naskar

机构 * University of Glasgow(格拉斯哥大学) Jadavpur University(贾瓦德普尔大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17642 2025-09-04 cs.CL cs.AI 57%

Banishing LLM Hallucinations Requires Rethinking Generalization

Johnny Li, Saksham Consul, Eda Zhou, James Wong, Naila Farooqui, Yuxin Ye, Nithyashree Manohar, Zhuxiaona Wei, Tian Wu, Ben Echols, Sharon Zhou, Gregory Diamos

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments I want to revisit some of the experiments in this paper, specifically figure 5

详情

展开后加载摘要…

URL PDF HTML 收藏