arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-01-27 至 2026-01-27 共收录 81 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 21 篇

2601.17588 2026-01-27 cs.AI cs.CL 79%

Intelligence Requires Grounding But Not Embodiment

智能需要具身但不需要身体

Marcus Ma, Shrikanth Narayanan

机构 * University of Southern California(南加州大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 本文提出智能需要基础性而非具身,通过定义智能的四个属性并论证非具身智能体可实现这些属性,从而得出结论:基础性是智能的必要条件。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06515 2026-01-27 cs.CV 79%

BigTokDetect: A Clinically-Informed Vision-Language Modeling Framework for Detecting Pro-Bigorexia Videos on TikTok

BigTokDetect: 一种临床指导的视觉-语言建模框架,用于检测TikTok上促进大肌肉畸形行为的视频

Minh Duc Chu, Kshitij Pawar, Zihao He, Roxanna Sharifi, Ross Sonnenblick, Magdalayna Curry, Laura D'Adamo, Lindsay Young, Stuart B Murray, Kristina Lerman

机构 * USC Information Sciences Institute(USC信息科学研究所) Keck School of Medicine, USC(USC凯克医学院) Department of Clinical Psychology, Drexel University(德雷塞尔大学临床心理学系) Department of Psychiatry and Biobehavioral Sciences, UCLA(UCLA精神病学与生物行为科学系)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

AI总结 BigTokDetect通过临床指导的视觉-语言模型,检测TikTok上促进大肌肉畸形行为的视频,建立了可扩展的有害内容缓解框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08905 2026-01-27 cs.AI cs.CL 79%

Grounding Synthetic Data Evaluations of Language Models in Unsupervised Document Corpora

在无监督文档语料中使语言模型的合成数据评估接地

Michael Majurski, Cynthia Matuszek

机构 * University of Maryland Baltimore County(马里兰大学巴尔的摩县分校)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 本研究提出了一种自动化方法,利用文档语料生成事实性合成数据评估,以提高语言模型在无监督文档中的评估效率和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18323 2026-01-27 cs.RO 78%

TC-IDM: Grounding Video Generation for Executable Zero-shot Robot Motion

TC-IDM:为可执行零样本机器人运动实现视频生成

Weishi Mi, Yong Bao, Xiaowei Chi, Xiaozhu Ju, Zhiyuan Qin, Kuangzhi Ge, Kai Tang, Peidong Jia, Shanghang Zhang, Jian Tang

机构 * Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心) State Key Laboratory of Multimedia Information Processing(多媒体信息处理国家重点实验室) School of Computer Science, Peking University(北京大学计算机科学学院)

专题命中 视觉定位与Grounding :grounding(title);vision-language model(abstract)

AI总结 TC-IDM通过工具中心逆动力学模型实现视频生成,提升机器人零样本任务的执行能力与泛化性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11031 2026-01-27 cs.LG cs.AI cs.CL 73%

Prefill-Guided Thinking for zero-shot detection of AI-generated images

预填充引导的思考用于AI生成图像的零样本检测

Zoher Kachwala, Danishjeet Singh, Danielle Yang, Filippo Menczer

机构 * Observatory on Social Media(社会媒体观察所) Indiana University(印第安纳大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.AI、cs.LG

AI总结 本文提出预填充引导思考方法,通过引导视觉-语言模型推理提升AI生成图像的零样本检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17844 2026-01-27 cs.HC cs.AI cs.LG 73%

RAICL: Retrieval-Augmented In-Context Learning for Vision-Language-Model Based EEG Seizure Detection

RAICL:基于视觉-语言模型的EEG癫痫检测的检索增强上下文学习

Siyang Li, Zhuoya Wang, Xiyan Gui, Xiaoqing Chen, Ziwei Wang, Yaozhi Wen, Dongrui Wu

机构 * Ministry of Education Key Laboratory of Image Processing and Intelligent Control, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(教育部图像处理与智能控制重点实验室,人工智能与自动化学院,华中科技大学) State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, Institute of Automation, Chinese Academy of Sciences(脑认知与脑启发智能技术国家重点实验室,自动化研究所,中国科学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.AI、cs.LG

AI总结 RAICL通过利用视觉-语言模型分析EEG波形图,实现了更高效的癫痫检测,无需重新训练,具有广泛临床应用前景。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11808 2026-01-27 cs.CV cs.AI cs.CL cs.CY cs.MM 73%

Labels or Input? Rethinking Augmentation in Multimodal Hate Detection

标签还是输入?重新思考多模态仇恨检测中的增强

Sahajpreet Singh, Kokil Jaidka, Subhayan Mukerjee

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

AI总结 本文提出通过提示优化、微调和自动化数据增强改进小型模型,开发多模态增强框架以提升隐含仇恨检测性能。

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00662 2026-01-27 cs.CV cs.CL cs.LG 73%

Mitigating the Modality Gap: Few-Shot Out-of-Distribution Detection with Multi-modal Prototypes and Image Bias Estimation

弥合模态差距:基于多模态原型和图像偏差估计的少样本分布外检测

Yimu Wang, Evelien Riddell, Adrian Chow, Sean Sedwards, Krzysztof Czarnecki

机构 * University of Waterloo(滑铁卢大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.LG

AI总结 本文提出SUPREME框架,通过引入多模态原型和图像偏差估计,有效缓解图像与文本之间的模态差距,提升少样本分布外检测性能。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17673 2026-01-27 cs.CV cs.AI 62%

Uni-RS: A Spatially Faithful Unified Understanding and Generation Model for Remote Sensing

Uni-RS: 一种用于遥感的具有空间忠实性的统一理解和生成模型

Weiyu Zhang, Yuan Hu, Yong Li, Yu Liu

机构 * Institute of Remote Sensing and Geographic Information System, School of Earth and Space Sciences, Peking University(遥感与地理信息系统研究所,地球与空间科学学院,北京大学) Department of Civil and Environmental Engineering, The Hong Kong University of Science and Technology(土木与环境工程系,香港科学与技术大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

AI总结 Uni-RS通过空间布局规划、空间感知查询监督和图像描述空间布局变化,提升遥感文本到图像生成的空间忠实性,同时保持多模态理解任务的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03906 2026-01-27 cs.CV 61%

From Filters to VLMs: Benchmarking Defogging Methods through Object Detection and Segmentation Performance

从滤波器到视觉语言模型:通过目标检测和分割性能评估去雾方法

Ardalan Aryashad, Parsa Razmara, Amin Mahjoub, Seyedarmin Azizi, Mahdi Salmani, Arad Firouzkouhi

机构 * University of Southern California(南加州大学)

专题命中 视觉定位与Grounding :visual language model(abstract);分类 cs.CV;VLM(comments)

AI总结 本文通过目标检测和分割性能评估,探讨了去雾方法在真实与合成环境中的有效性,揭示了视觉语言模型在恶劣天气下的应用潜力。

Comments Accepted at WACV 2026 Proceedings (Oral), 5th Workshop on Image, Video, and Audio Quality Assessment in Computer Vision, with a focus on VLM and Diffusion Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18076 2026-01-27 cs.LG 57%

Comparison requires valid measurement: Rethinking attack success rate comparisons in AI red teaming

比较需要有效的测量:重新思考AI红队中的攻击成功率比较

Alexandra Chouldechova, A. Feder Cooper, Solon Barocas, Abhinav Palia, Dan Vann, Hanna Wallach

机构 * Microsoft Research(微软研究院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

AI总结 本文重新审视AI红队中攻击成功率的比较,指出其有效性问题并提出测量理论和统计学方法以评估比较的合理性。

Journal ref NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17435 2026-01-27 cs.SE cs.AI 57%

Towards a Declarative Agentic Layer for Intelligent Agents in MCP-Based Server Ecosystems

迈向基于MCP服务器生态系统的声明性代理层

Maria Jesus Rodriguez-Sanchez, Manuel Noguera, Angel Ruiz-Zafra, Kawtar Benghazi

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出了一种声明性代理层,旨在解决基于MCP服务器生态系统中代理工作流的可靠性问题,通过明确的架构结构提升任务执行的可验证性和可重复性。

Comments 12 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08710 2026-01-27 cs.CV 57%

SceneSplat++: A Large Dataset and Comprehensive Benchmark for Language Gaussian Splatting

SceneSplat++: 一个大规模数据集和全面的基准用于语言高斯点云

Mengjiao Ma, Qi Ma, Yue Li, Jiahuan Cheng, Runyi Yang, Bin Ren, Nikola Popovic, Mingqiang Wei, Nicu Sebe, Luc Van Gool, Theo Gevers, Martin R. Oswald, Danda Pani Paudel

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学) ETH Zürich(苏黎世联邦理工学院) University of Amsterdam(阿姆斯特丹大学) Johns Hopkins University(约翰·霍普金斯大学) University of Pisa(比萨大学) University of Trento(特伦托大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 SceneSplat++提出一个大规模数据集和基准,评估语言高斯点云方法,展示可推广方法在3D理解中的优势。

Comments 15 pages, codes, data and benchmark are released at https://scenesplatpp.gaussianworld.ai/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18267 2026-01-27 cs.IR 50%

Orchestrating Specialized Agents for Trustworthy Enterprise RAG

协调专用代理以实现可信的企业RAG

Xincheng You, Qi Sun, Neha Bora, Huayi Li, Shubham Goel, Kang Li, Sean Culatana

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 ADORE通过结构化记忆库和迭代协调机制,提升企业RAG在高风险决策中的可追溯性和证据完整性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17219 2026-01-27 cs.RO 50%

Advancing Improvisation in Human-Robot Construction Collaboration: Taxonomy and Research Roadmap

推动人机协同建造中的即兴能力:分类与研究路线图

David Wireko Atibila, Vineet R. Kamat, Carol C. Menassa

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文提出六级分类法,分析人机协同建造中即兴能力的发展现状与未来研究方向,指出技术、概念和方法学三方面障碍,建议通过增强现实、大语言模型和云系统提升人机协作水平。

Comments 73 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17002 2026-01-27 cs.CL 50%

RAM-SD: Retrieval-Augmented Multi-agent framework for Sarcasm Detection

RAM-SD:基于检索的多智能体框架用于讽刺检测

Ziyang Zhou, Ziqi Liu, Yan Wang, Yiming Lin, Yangbin Chen

机构 * Xi’an Jiaotong–Liverpool University(西安交通大学利物浦大学)

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 RAM-SD通过多智能体框架提升讽刺检测性能,实现77.74%的宏F1值,提供透明推理轨迹。

Comments 12 pages, 4 figures, 6 tables, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 文档图表理解 1 篇

2601.18735 2026-01-27 cs.AI cs.LG 62%

Why Keep Your Doubts to Yourself? Trading Visual Uncertainties in Multi-Agent Bandit Systems

为何保留你的怀疑?多智能体老虎机系统中的视觉不确定性交易

Jusheng Zhang, Yijia Fan, Kaitong Cai, Jing Yang, Jiawei Yao, Jian Wang, Guanlong Qu, Ziliang Chen, Keze Wang

机构 * Sun Yat-sen University(中山大学) University of Washington(华盛顿大学) Snap Inc.(Snap公司) Syracuse University(雪城大学)

专题命中 文档图表理解 :vision-language model(abstract);分类 cs.AI、cs.LG

AI总结 Agora通过去中心化市场交易机制提升多智能体系统在视觉任务中的协调效率与经济性。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

3. GUI与屏幕智能体 4 篇

2503.18712 2026-01-27 cs.CV 70%

LLaVAction: evaluating and training multi-modal large language models for action understanding

LLaVAction:评估和训练多模态大语言模型进行动作理解

Haozhe Qi, Shaokai Ye, Alexander Mathis, Mackenzie W. Mathis

专题命中 GUI与屏幕智能体 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

AI总结 LLaVAction通过引入动作标记和两阶段流程提升多模态大语言模型的动作理解能力,显著提升基准测试性能。

Comments https://github.com/AdaptiveMotorControlLab/LLaVAction

Journal ref International Conference on Learning Representations (ICLR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10046 2026-01-27 cs.AI 70%

SimWorld-Robotics: Synthesizing Photorealistic and Dynamic Urban Environments for Multimodal Robot Navigation and Collaboration

SimWorld-Robotics: 为多模态机器人导航与协作合成逼真动态城市环境

Yan Zhuang, Jiawei Ren, Xiaokang Ye, Jianzhi Shen, Ruixuan Zhang, Tianai Yue, Muhammad Faayez, Xuhong He, Ziqiao Ma, Lianhui Qin, Zhiting Hu, Tianmin Shu

机构 * University of Virginia(弗吉尼亚大学) UC San Diego(加州大学圣地亚哥分校) Johns Hopkins University(约翰霍普金斯大学) Carnegie Mellon University(卡内基梅隆大学) University of Michigan(密歇根大学)

专题命中 GUI与屏幕智能体 :vision-language model(abstract);grounding(abstract);分类 cs.AI

AI总结 SimWorld-Robotics通过合成逼真动态城市环境,提出两个多模态机器人基准测试,评估机器人在复杂场景中的导航、协作与通信能力。

Comments Conference: NeurIPS 2025 (main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17885 2026-01-27 cs.CV cs.AI cs.RO 62%

PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation

PEAfowl:感知增强的多视角视觉-语言-动作用于双臂操作

Qingyu Fan, Zhaoxiang Li, Yi Lu, Wang Chen, Qiu Shen, Xiao-xiao Long, Yinghao Cai, Tao Lu, Shuo Wang, Xun Cao

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Nanjing University(南京大学)

专题命中 GUI与屏幕智能体 :grounding(abstract);分类 cs.CV、cs.AI

AI总结 PEAfowl通过增强感知的多视角视觉-语言-动作策略,提升双臂操作在复杂环境中的稳定性和成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17722 2026-01-27 cs.AI 57%

EntWorld: A Holistic Environment and Benchmark for Verifiable Enterprise GUI Agents

EntWorld: 一个综合环境和基准,用于可验证的企业GUI代理

Ying Mo, Yu Bai, Dapeng Sun, Yuqian Shi, Yukai Miao, Li Chen, Dan Li

机构 * Zhongguancun Laboratory(中关村实验室) Tsinghua University(清华大学)

专题命中 GUI与屏幕智能体 :multimodal large language model(abstract);分类 cs.AI

AI总结 EntWorld是一个综合的企业GUI代理基准,通过基于模式的任务生成和SQL验证机制,评估代理在企业环境中的表现,揭示当前代理能力的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与鲁棒性 10 篇

2510.26441 2026-01-27 cs.CV 85%

A-TPT: Angular Diversity Calibration Properties for Test-Time Prompt Tuning of Vision-Language Models

A-TPT:面向视觉语言模型测试时提示微调的角多样性校准特性

Shihab Aaqil Ahamed, Udaya S. K. P. Miriya Thanthrige, Ranga Rodrigo, Muhammad Haris Khan

机构 * Dept. of Electronic and Telecommunication Engineering, University of Moratuwa(摩图瓦大学电子与电信工程系) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(abstract);grounding(abstract);分类 cs.CV

AI总结 A-TPT通过引入角度多样性提升视觉语言模型测试时提示微调的校准性能,有效减少校准误差并提升适应能力。

Comments Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17082 2026-01-27 cs.CY cs.AI cs.CL cs.CV cs.LG 82%

Do VLMs Have a Moral Backbone? A Study on the Fragile Morality of Vision-Language Models

视觉-语言模型是否具有道德基础?对视觉语言模型脆弱道德的研究

Zhining Liu, Tianyi Wang, Xiao Lin, Penghao Ouyang, Gaotang Li, Ze Yang, Hui Liu, Sumit Keswani, Vishwa Pardeshi, Huijun Zhao, Wei Fan, Hanghang Tong

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Amazon(亚马逊) Fidelity Investments(富达投资)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 研究发现视觉语言模型在面对文本和视觉扰动时道德判断易变,需通过轻量干预提升道德鲁棒性以确保负责任的部署。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11908 2026-01-27 cs.RO 75%

Safe Learning for Contact-Rich Robot Tasks: A Survey from Classical Learning-Based Methods to Safe Foundation Models

接触丰富机器人任务的安全学习:从经典学习方法到安全基础模型的综述

Heng Zhang, Rui Dai, Gokhan Solak, Pokuang Zhou, Yu She, Arash Ajoudani

机构 * Human-Robot Interfaces and Interaction Lab, Istituto Italiano di Tecnologia, Genova, Italy(人类-机器人接口与交互实验室,意大利技术研究院,热那亚,意大利) Ph.D. program of national interest in Robotics and Intelligent Machines (DRIM) and Università di Genova, Genoa, Italy(机器人与智能机器国家利益博士项目(DRIM)和热那亚大学,热那亚,意大利) Edwardson School of Industrial Engineering, Purdue University, West Lafayette, IN, USA(工业工程埃德华森学校,普渡大学,西拉法伊斯,美国)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);VLM(abstract);grounding(abstract)

AI总结 本文综述了接触丰富机器人任务的安全学习方法,探讨了从经典学习方法到安全基础模型的发展,分析了安全探索与执行的关键技术及未来方向。

Comments version 2

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13379 2026-01-27 cs.AI cs.CV 62%

The Art of Saying "Maybe": A Conformal Lens for Uncertainty Benchmarking in VLMs

也许的艺术:一种用于VLMs不确定性基准测试的共形透镜

Asif Azad, Mohammad Sadat Hossain, MD Sadik Hossain Shanto, M Saifur Rahman, Md Rizwan Parvez

机构 * Bangladesh University of Engineering and Technology(孟加拉工程与技术大学) Qatar Computing Research Institute(卡塔尔计算研究所)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种用于VLMs不确定性基准测试的共形透镜,评估了18种最先进的VLMs在6个多元数据集上的表现,发现更大模型在不确定性量化方面表现更佳,而数学和推理任务则表现出较差的不确定性性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17357 2026-01-27 cs.LG cs.AI 62%

Spectral Geometry for Deep Learning: Compression and Hallucination Detection via Random Matrix Theory

深度学习的谱几何:通过随机矩阵理论实现压缩与幻觉检测

Davide Ettori

机构 * Politecnico di Milano(米兰理工学院) University of Illinois Chicago(伊利诺伊大学芝加哥分校)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出基于谱几何和随机矩阵理论的统一框架,通过分析隐藏激活的特征值结构,实现幻觉检测与模型压缩,提升深度学习的可靠性和效率。

Comments Master thesis, MS in Computer Science, University of Illinois Chicago, defended November 21, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17093 2026-01-27 cs.LG cs.AI 62%

The Triangle of Similarity: A Multi-Faceted Framework for Comparing Neural Network Representations

相似性的三角形:一种多维框架用于比较神经网络表示

Olha Sirikova, Alvin Chan

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出相似性三角形框架,通过静态、功能和稀疏性三个视角比较神经网络表示,揭示架构对表示相似性的影响及剪枝对模型核心的影响。

Comments Accepted to AAAI 2026 Workshop on AI for Scientific Research (AI4Research)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18698 2026-01-27 cs.CV 57%

Are Video Generation Models Geographically Fair? An Attraction-Centric Evaluation of Global Visual Knowledge

视频生成模型在地理上是否公平?一种以吸引力为中心的全球视觉知识评估

Xiao Liu, Jiawei Zhang

机构 * IFM Lab, Department of Computer Science University of California, Davis(信息融合实验室,计算机科学系加州大学戴维斯分校)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

AI总结 本文通过GAP框架评估了文本到视频模型在地理公平性上的表现,发现模型在不同地区和文化群体中具有相对均匀的地理相关视觉知识,表明其在全球应用中的潜力。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21885 2026-01-27 cs.CV cs.MM cs.RO 57%

Integrating Multi-Modal Sensors: A Review of Fusion Techniques for Intelligent Vehicles

多模态传感器整合:智能车辆融合技术综述

Chuheng Wei, Ziye Qin, Ziyan Zhang, Guoyuan Wu, Matthew J. Barth

机构 * College of Engineering, Center for Environmental Research and Technology, University of California at Riverside(工程学院、环境研究与技术中心、加州大学河滨分校) School of Transportation and Logistics, Southwest Jiaotong University(交通运输与物流学院、西南交通大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

AI总结 本文综述了多传感器融合技术在自动驾驶中的应用,分析了深度学习方法、多模态数据集及新兴趋势,强调了其提升系统适应性和鲁棒性的潜力。

Comments Accepted by IEEE IV 2025

Journal ref Proceedings of the 2025 IEEE Intelligent Vehicles Symposium (IV), pp. 1817-1824, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.03473 2026-01-27 cs.CY cs.AI cs.CL 57%

Exploring LGBTQ+ Bias in Generative AI Answers across Different Country and Religious Contexts

探索不同国家和宗教背景下生成AI回答中的LGBTQ+偏见

Lilla Vicsek, Anna Vancsó, Mike Zajko, Judit Takacs

专题命中 幻觉与鲁棒性 :grounding(abstract);分类 cs.AI

AI总结 本研究探讨生成AI在不同国家和宗教背景下对LGBTQ+偏见的回应,发现ChatGPT 3.5表现出文化相对主义,而Bard强调人权并提供更多支持。

Comments Replacement version -- includes link to BD&S journal publication (significantly revised) in abstract, but the manuscript here remains unchanged from the original arXiv version

详情

展开后加载摘要…

URL PDF HTML 收藏