arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7409 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7409 篇

2510.21679 2025-10-27 cs.AI 57%

A Multimodal Benchmark for Framing of Oil & Gas Advertising and Potential Greenwashing Detection

Gaku Morio, Harri Rowlands, Dominik Stammbach, Christopher D. Manning, Peter Henderson

机构 * Hitachi, Ltd.(日立公司) Stanford University(斯坦福大学) Centre for the Acceleration of Social Technology(社会技术加速中心) Princeton University(普林斯顿大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.AI

Comments Forthcoming in NeurIPS 2025 Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21449 2025-10-27 cs.CV 57%

MoniTor: Exploiting Large Language Models with Instruction for Online Video Anomaly Detection

Shengtian Yang, Yue Feng, Yingshi Liu, Jingrou Zhang, Jie Qin

机构 * College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(人工智能学院,南京航空航天大学) Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education, China(脑机智能技术重点实验室,教育部,中国)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025. The first two authors hold equal contributions

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21302 2025-10-27 cs.AI cs.RO 57%

Towards Reliable Code-as-Policies: A Neuro-Symbolic Framework for Embodied Task Planning

Sanghyun Ahn, Wonje Choi, Junyong Lee, Jinwoo Park, Honguk Woo

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted at NeurIPS 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20279 2025-10-27 cs.LG 57%

ResearchGPT: Benchmarking and Training LLMs for End-to-End Computer Science Research Workflows

Penghao Wang, Yuhao Zhou, Mengxuan Wu, Ziheng Qin, Bangyuan Zhu, Shengbin Huang, Xuanlei Zhao, Panpan Zhang, Xiaojiang Peng, Yuzhang Shang, Jianfei Yang, Zheng Zhu, Tianlong Chen, Zhangyang Wang, Kai Wang

机构 * NUS(国立新加坡大学) NTU(国立科技大学) SZTU(深圳技术大学) UCF(佛罗里达大学) GigaAI(GigaAI研究所) UNC(北卡罗来纳大学教堂山分校) UT Austin(得克萨斯大学奥斯汀分校)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06958 2025-10-27 cs.CY cs.AI cs.MA 57%

Simulating Society Requires Simulating Thought

Chance Jiajie Li, Jiayi Wu, Zhenze Mo, Ao Qu, Yuhan Tang, Kaiya Ivy Zhao, Yulu Gan, Jie Fan, Jiangbo Yu, Jinhua Zhao, Paul Liang, Luis Alonso, Kent Larson

机构 * MIT Media Lab(麻省理工学院媒体实验室) MIT EECS(麻省理工学院电子工程与计算机科学系) MIT BCS(麻省理工学院生物工程与计算机科学系) MIT IDSS(麻省理工学院国际设计与研究系统) MIT CEE(麻省理工学院土木与环境工程系) MIT DUSP(麻省理工学院设计与科学计划) MIT Architecture(麻省理工学院建筑系) Northeastern University(东北大学) Brown University(布朗大学) McGill University(麦吉尔大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments NeurIPS 2025 (Position Paper Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18812 2025-10-27 cs.CV 57%

SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models

Ye Sun, Hao Zhang, Henghui Ding, Tiehua Zhang, Xingjun Ma, Yu-Gang Jiang

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03325 2025-10-27 cs.CL cs.AI 57%

Electronic Circuit Principles of Large Language Models

Qiguang Chen, Libo Qin, Jinhao Liu, Dengyun Peng, Jiaqi Wang, Mengkang Hu, Zhi Chen, Wanxiang Che, Ting Liu

机构 * Harbin Institute of Technology(哈尔滨工业大学) Central South University(中南大学) The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) ByteDance Seed (China)(字节跳动种子(中国))

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19574 2025-10-23 cs.CV cs.CR 57%

Can You Trust What You See? Alpha Channel No-Box Attacks on Video Object Detection

Ariana Yi, Ce Zhou, Liyang Xiao, Qiben Yan

机构 * Mission San Jose High School(Mission San Jose 高中) Missouri University of Science and Technology(密苏里科学与技术大学) Michigan State University(密歇根州立大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18476 2025-10-22 cs.AI cs.CL 57%

Probabilistic Modeling of Intentions in Socially Intelligent LLM Agents

Feifan Xia, Yuyang Fang, Defang Li, Yantong Xie, Weikang Li, Yang Li, Deguo Xia, Jizhou Huang

机构 * Baidu Inc(百度公司) Imperial College London(伦敦帝国学院) Zhejiang University(浙江大学) Carnegie Mellon University(卡内基梅隆大学) Peking University(北京大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11741 2025-10-22 cs.AI cs.CR 57%

MTRE: Multi-Token Reliability Estimation for Hallucination Detection in VLMs

Geigh Zollicoffer, Minh Vu, Manish Bhattarai

机构 * Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03621 2025-10-21 cs.CV 57%

DynVFX: Augmenting Real Videos with Dynamic Content

Danah Yatim, Rafail Fridman, Omer Bar-Tal, Tali Dekel

机构 * Weizmann Institute of Science(魏兹曼科学研究所)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Project page: https://dynvfx.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16530 2025-10-21 cs.LG stat.ML 57%

Realizing LLMs' Causal Potential Requires Science-Grounded, Novel Benchmarks

Ashutosh Srivastava, Lokesh Nagalapatti, Gautam Jajoo, Aniket Vashishtha, Parameswari Krishnamurthy, Amit Sharma

机构 * Dept. of Computer Science IIIT Hyderabad(IIIT Hyderabad 计算机科学系) Dept. of Computer Science IIT Bombay(IIT Bombay 计算机科学系) Dept. of Computer Science BITS Pilani(BITS Pilani 计算机科学系) Dept. of Computer Science UIUC(UIUC 计算机科学系) Microsoft Research Bengaluru, India(微软研究院(印度班加罗尔))

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16374 2025-10-21 cs.AI 57%

Before you <think>, monitor: Implementing Flavell's metacognitive framework in LLMs

Nick Oh

机构 * socius labs(socius实验室)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Presented at the Workshop on the Application of LLM Explainability to Reasoning and Planning at COLM 2025 (non-archival)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15895 2025-10-21 cs.HC cs.AI cs.SD 57%

BREATH: A Bio-Radar Embodied Agent for Tonal and Human-Aware Diffusion Music Generation

Yunzhe Wang, Xinyu Tang, Zhixun Huang, Xiaolong Yue, Yuxin Zeng

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted by LLM4Music @ ISMIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04232 2025-10-20 cs.LG 57%

Rethinking Layer-wise Gaussian Noise Injection: Bridging Implicit Objectives and Privacy Budget Allocation

Qifeng Tan, Shusen Yang, Xuebin Ren, Yikai Zhang

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments Errors were found in the experimental data preprocessing, which affected the reported results and conclusions. The paper is being revised and a corrected version will be resubmitted

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14944 2025-10-17 cs.CL cs.AI cs.CE 57%

MetaBench: A Multi-task Benchmark for Assessing LLMs in Metabolomics

Yuxing Lu, Xukai Zhao, J. Ben Tamo, Micky C. Nnamdi, Rui Peng, Shuang Zeng, Xingyu Hu, Jinzhuo Wang, May D. Wang

机构 * Wallace H. Coulter Department of Biomedical Engineering, Georgia Institute of Technology and Emory University(沃克生物医学工程部门,佐治亚理工学院和埃默里大学) College of Future of Technology, Peking University(未来技术学院,北京大学) School of Architecture, Tsinghua University(建筑学院,清华大学) School of Electrical and Computer Engineering, Georgia Institute of Technology(电气与计算机工程学院,佐治亚理工学院) School of Computer Science, Georgia Institute of Technology(计算机科学学院,佐治亚理工学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 22 pages, 6 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14763 2025-10-17 cs.CL cs.AI 57%

COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes

Yunwen Li, Shuangshuang Ying, Xingwei Qu, Xin Li, Sheng Jin, Minghao Liu, Zhoufutu Wen, Tianyu Zheng, Xeron Du, Qiguang Chen, Jiajun Shi, Wangchunshu Zhou, Jiazhan Feng, Wanjun Zhong, Libo Qin, Stephen Huang, Wanxiang Che, Chenghua Lin, Eli Zhang

机构 * M-A-P 2077AI

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12217 2025-10-17 cs.CL cs.AI 57%

HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment

Ali Mekky, Omar El Herraoui, Preslav Nakov, Yuxia Wang

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10300 2025-10-16 cs.CC cs.AI cs.IT cs.SY eess.SY math.IT q-bio.NC 57%

The Algorithmic Regulator

Giulio Ruffini

机构 * Neuroelectrics(神经电医学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 2 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02527 2025-10-16 cs.CL cs.LG 57%

I Have No Mouth, and I Must Rhyme: Uncovering Internal Phonetic Representations in LLaMA 3.2

Oliver McLaughlin, Arjun Khurana, Jack Merullo

机构 * Brown University(布朗大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12763 2025-10-15 eess.SP cs.AI q-bio.QM 57%

Disentangling Neurodegeneration with Brain Age Gap Prediction Models: A Graph Signal Processing Perspective

Saurabh Sihag, Gonzalo Mateos, Alejandro Ribeiro

机构 * Department of Electrical and Computer Engineering at the University at Albany, SUNY(纽约州立大学阿尔巴尼分校电气与计算机工程系) Department of Electrical and Computer Engineering at the University of Rochester(罗切斯特大学电气与计算机工程系) Department of Electrical and Systems Engineering at the University of Pennsylvania(宾夕法尼亚大学电气与系统工程系)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted for publication in IEEE Signal Processing Magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11602 2025-10-14 cs.CL cs.LG 57%

Deconstructing Attention: Investigating Design Principles for Effective Language Modeling

Huiyin Xue, Nafise Sadat Moosavi, Nikolaos Aletras

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11063 2025-10-14 cs.CV 57%

LSVOS 2025 Challenge Report: Recent Advances in Complex Video Object Segmentation

Chang Liu, Henghui Ding, Kaining Ying, Lingyi Hong, Ning Xu, Linjie Yang, Yuchen Fan, Mingqi Gao, Jingkun Chen, Yunqi Miao, Gengshen Wu, Zhijin Qin, Jungong Han, Zhixiong Zhang, Shuangrui Ding, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Jiaqi Wang, Chang Soo Lim, Joonyoung Moon, Donghyeon Cho, Tingmin Li, Yixuan Li, Yang Yang, An Yan, Leilei Cao, Feng Lu, Ran Hong, Youhai Jiang, Fengjie Zhu, Yujie Xie, Hongyang Zhang, Zhihui Liu, Shihai Ruan, Quanzhu Niu, Dengxian Gong, Shihao Chen, Tao Zhang, Yikang Zhou, Haobo Yuan, Lu Qi, Xiangtai Li, Shunping Ji, Ran Hong, Feng Lu, Leilei Cao, An Yan, Alexey Nekrasov, Ali Athar, Daan de Geus, Alexander Hermans, Bastian Leibe

专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV

Comments 16 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09421 2025-10-13 cs.CL cs.AI 57%

On the Representations of Entities in Auto-regressive Large Language Models

Victor Morand, Josiane Mothe, Benjamin Piwowarski

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted at BlackBoxNLP@EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09274 2025-10-13 cs.CV 57%

MomentSeg: Moment-Centric Sampling for Enhanced Video Pixel Understanding

Ming Dai, Sen Yang, Boqiang Duan, Wankou Yang, Jingdong Wang

机构 * Southeast University(东南大学) Baidu VIS(百度视觉)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08849 2025-10-13 cs.CV 57%

FOLK: Fast Open-Vocabulary 3D Instance Segmentation via Label-guided Knowledge Distillation

Hongrui Wu, Zhicheng Gao, Jin Cao, Kelu Yao, Wen Shen, Zhihua Wei

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21787 2025-10-13 cs.CV cs.CL 57%

DeHate: A Stable Diffusion-based Multimodal Approach to Mitigate Hate Speech in Images

Dwip Dalal, Gautam Vashishtha, Anku Rani, Aishwarya Reganti, Parth Patwa, Mohd Sarique, Chandan Gupta, Keshav Nath, Viswanatha Reddy, Vinija Jain, Aman Chadha, Amitava Das, Amit Sheth, Asif Ekbal

机构 * MIT Media Lab, USA(麻省理工学院媒体实验室) Stanford University, USA(斯坦福大学) Amazon GenAI, USA(亚马逊生成人工智能) University of South Carolina, USA(南卡罗来纳大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Defactify 3 workshop at AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19114 2025-10-13 cs.CL cs.IR cs.LG 57%

Understanding and Improving Information Preservation in Prompt Compression for LLMs

Weronika Łajewska, Momchil Hardalov, Laura Aina, Neha Anna John, Hang Su, Lluís Màrquez

机构 * University of Stavanger(斯塔万格大学) AWS AI Labs(AWS AI实验室) Technical University of Catalonia (UPC)(加泰罗尼亚理工大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments Accepted to EMNLP 2025 (Findings), 22 pages, 6 figures, 24 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16456 2025-10-10 q-bio.NC cs.CV 57%

Language learning shapes visual category-selectivity in deep neural networks

Zitong Lu, Yuxin Wang

机构 * MIT McGovern Institute for Brain Research(麻省理工学院麦戈文脑研究所) University of Cincinnati College of Medicine(辛辛那提大学医学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07000 2025-10-09 cs.CL cs.AI 57%

Pragyaan: Designing and Curating High-Quality Cultural Post-Training Datasets for Indian Languages

Neel Prabhanjan Rachamalla, Aravind Konakalla, Gautam Rajeev, Ashish Kulkarni, Chandra Khatri, Shubham Agarwal

机构 * Krutrim AI

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏