arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7464 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7464 篇

2501.13795 2026-05-19 cs.CV 74%

Training-Free Zero-Shot Temporal Action Detection with Vision-Language Models

无需训练的零样本时序动作检测与视觉-语言模型

Chaolei Han, Hongsong Wang, Jidong Kuang, Lei Zhang, Jie Gui

机构 * Southeast University School of Cyber Science and Engineering(东南大学网络安全科学与工程学院) Southeast University School of Computer Science and Engineering(东南大学计算机科学与工程学院) Nanjing Normal University School of Electrical Engineering and Automation(南京师范大学电气工程与自动化学院)

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

AI总结 本文提出一种无需训练的零样本时序动作检测方法FreeZAD,利用现有的视觉-语言模型直接对未标记视频中的未知活动进行分类和定位,无需额外微调或适应,并通过LogOIC和频率基于的动作校准以及测试时适应策略提升性能。

Journal ref IEEE Transactions on Multimedia, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07692 2026-04-30 cs.LG 74%

Tree-of-Evidence: Efficient "System 2" Search for Faithful Multimodal Grounding

证据树:高效的'系统2'搜索用于忠实的多模态接地

Micky C. Nnamdi, Benoit L. Marteau, Yishan Zhong, J. Ben Tamo, May D. Wang

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.LG

AI总结 本文提出Tree-of-Evidence算法,通过离散优化问题实现多模态模型的可解释性,通过证据瓶颈和束搜索生成可审计的证据集,保留高预测性能并提升决策一致性。

Journal ref ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23588 2026-04-28 cs.AI cs.CL cs.IR 74%

FinGround: Detecting and Grounding Financial Hallucinations via Atomic Claim Verification

FinGround:通过原子性声明验证检测和支撑金融幻觉

Dongxin Guo, Jikun Wu, Siu Ming Yiu

机构 * The University of Hong Kong(香港大学) Stellaris AI Limited(Stellaris AI有限公司)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

AI总结 FinGround通过三阶段验证-支撑流程,针对金融文档问答中的计算错误进行验证和修正,有效降低幻觉率,其在金融领域具有重要应用价值。

Comments Accepted to ACL 2026 Industry Track. 14 pages, 1 figure, 14 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20246 2026-04-23 cs.RO cs.AI 74%

Cortex 2.0: Grounding World Models in Real-World Industrial Deployment

Cortex 2.0:在真实工业部署中将世界模型接地

Adriana Aida, Walida Amer, Katarina Bankovic, Dhruv Behl, Fabian Busch, Annie Bhalla, Minh Duong, Florian Gienger, Rohan Godse, Denis Grachev, Ralf Gulde, Elisa Hagensieker, Junpeng Hu, Shivam Joshi, Tobias Knoblauch, Likith Kumar, Damien LaRocque, Keerthana Lokesh, Omar Moured, Khiem Nguyen, Christian Preyss, Ranjith Sriganesan, Vikram Singh, Carsten Sponner, Anh Tong, Dominik Tuscher, Marc Tuscher, Pavan Upputuri

机构 * Sereact GmbH(Sereact公司)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

AI总结 Cortex 2.0通过在视觉潜在空间中生成候选未来轨迹并评分,实现了从反应式控制到计划与行动的转变,展示了在复杂工业环境中可靠的世界模型规划。

Comments 20 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08226 2026-04-10 cs.AI cs.HC cs.SY eess.SY 74%

Grounding Clinical AI Competency in Human Cognition Through the Clinical World Model and Skill-Mix Framework

通过临床世界模型和技能混合框架在人类认知中奠定临床AI能力

Seyed Amir Ahmad Safavi-Naini, Elahe Meftah, Josh Mohess, Pooya Mohammadi Kazaj, Georgios Siontis, Zahra Atf, Peter R. Lewis, Mauricio Reyes, Girish Nadkarni, Roland Wiest, Stephan Windecker, Christoph Grani, Ali Soroush, Isaac Shiri

机构 * Department of Cardiology, Inselspital, Bern University Hospital, University of Bern(伯尔尼大学医院心脏病学系,伯尔尼大学) Department of Digital Medicine, Bern University Hospital, University of Bern(伯尔尼大学医院数字医学系,伯尔尼大学) Division of Data-Driven and Digital Medicine (D3M), Icahn School of Medicine at Mount Sinai(西奈山伊坎医学院数据驱动与数字医学部) Clinical Research Development Center, Amir Oncology Teaching Hospital, Shiraz University of Medical Sciences(设拉子医科大学阿米尔肿瘤教学医院临床研究发展中心) Graduate School for Cellular and Biomedical Sciences, University of Bern(伯尔尼大学细胞与生物医学研究生院) Faculty of Business and Information Technology, Ontario Tech University(安大略理工大学商业与信息技术学院) Department of Radiation Oncology, Inselspital, Bern University Hospital and University of Bern(伯尔尼大学医院放射肿瘤学系,伯尔尼大学) ARTORG Center for Biomedical Engineering Research, University of Bern(伯尔尼大学ARTORG生物医学工程研究中心) The Charles Bronfman Institute of Personalised Medicine, Icahn School of Medicine at Mount Sinai(西奈山伊坎医学院查尔斯·布朗夫曼个性化医学研究所) University Institute of Diagnostic and Interventional Neuroradiology, Inselspital, Bern University Hospital, University of Bern(伯尔尼大学医院诊断与介入神经放射学大学研究所,伯尔尼大学) Translational Imaging Center (TIC), Swiss Institute for Translational and Entrepreneurial Medicine(瑞士转化与创业医学研究所转化影像中心) Henry D. Janowitz Division of Gastroenterology, Icahn School of Medicine at Mount Sinai(西奈山伊坎医学院亨利·D·雅诺维茨消化内科)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

AI总结 本文提出临床世界模型和技能混合框架,通过八维定义临床能力空间,为AI在医疗场景中的能力评估和验证提供结构化方法。

Comments Code, data (Clinical AI Skill-Mix dimension specifications), and an exploratory dashboard are available at https://github.com/Sdamirsa/Clinical-World-Model

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08033 2026-04-10 cs.AI cs.MA cs.NI 74%

IoT-Brain: Grounding LLMs for Semantic-Spatial Sensor Scheduling

IoT-Brain:为语义-空间传感器调度 grounding LLMs

Zhaomeng Zhou, Lan Zhang, Junyang Wang, Mu Yuan, Junda Lin, Jinke Song

机构 * University of Science and Technology of China(中国科学技术大学) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院) The Chinese University of Hong Kong(香港中文大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

AI总结 本文提出IoT-Brain系统,通过引入空间轨迹图(STG)解决LLM在语义-空间传感器调度中的表示、推理和优化缺陷,提升任务成功率37.6%并显著降低资源消耗。

Comments To appear in ACM MobiCom 2026; 13 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00503 2026-04-08 cs.CV 74%

PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training

PET-DINO:将视觉线索统一到Grounding DINO中通过提示增强训练

Weifu Fu, Jinyang Li, Bin-Bin Gao, Jialin Li, Yuhuan Lin, Hanqiu Deng, Wenbing Tao, Yong Liu, Chengjie Wang

机构 * YouTu Lab, Tencent(腾讯优图实验室) Huazhong University of Science and Technology(华中科技大学) Kling Team, Kuaishou Technology(快手科技可灵团队)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

AI总结 PET-DINO通过提示增强训练策略,统一视觉线索到Grounding DINO中,提升零样本目标检测性能。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11782 2026-04-07 cs.HC cs.AI 74%

Human-AI Collaborative Game Testing with Vision Language Models

人机协作的游戏测试与视觉语言模型

Boran Zhang, Muhan Xu, Zhijun Pan

机构 * Zhengzhou University(郑州大学) University of the Arts London(伦敦艺术大学) Royal College of Art(皇家艺术学院)

专题命中 视觉定位与Grounding :vision language model(title);分类 cs.AI

AI总结 本文研究了如何利用AI提升游戏测试效率,通过800个测试用例和276名参与者实验,发现AI辅助能显著提高缺陷识别,但AI错误会影响人类决策,强调优化人机协作的重要性。

Comments Experiment Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26584 2026-04-06 cs.CV 74%

Scene Grounding In the Wild

野外场景的场景定位

Tamir Cohen, Leo Segre, Shay Shomer-Chai, Shai Avidan, Hadar Averbuch-Elor

机构 * Tel Aviv University(特拉维夫大学) Cornell University(康奈尔大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

AI总结 本文提出一种框架,通过将部分重建与完整的参考模型对齐,提升无视觉重叠时的大尺度场景重建一致性。方法基于伪合成渲染,利用3D高斯点云和逆特征优化,引入WikiEarth数据集验证效果。

Comments CVPR 2026. Project page at https://tau-vailab.github.io/SceneGround/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15154 2026-03-18 eess.IV cs.CV 74%

Vision-Language Model Based Multi-Expert Fusion for CT Image Classification

基于视觉-语言模型的多专家融合用于CT图像分类

Jianfa Bai, Kejin Lu, Runtian Yuan, Qingqiu Li, Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng

机构 * College of Computer Science and Artificial Intelligence, Shanghai Key Laboratory of Intelligent Information Processing, Fudan University(复旦大学计算机科学与人工智能学院,上海智能信息处理重点实验室) University of Oxford(牛津大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

AI总结 本文提出一种三阶段源感知多专家框架,通过构建肺部感知3D专家、开发MedSigLIP基专家和训练源分类器,提升多源CT图像中新冠检测的鲁棒性,实验结果显示在不同阶段模型在宏F1、ACC和AUC指标上均取得优异成绩。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09036 2026-03-11 cs.LG 74%

SCALAR: Learning and Composing Skills through LLM Guided Symbolic Planning and Deep RL Grounding

SCALAR:通过LLM引导的符号规划和深度RL接地学习与组合技能

Renos Zabounidis, Yue Wu, Simon Stepputtis, Woojun Kim, Yuanzhi Li, Tom Mitchell, Katia Sycara

专题命中 视觉定位与Grounding :grounding(title);分类 cs.LG

AI总结 SCALAR通过结合LLM引导的符号规划与深度RL,实现技能学习与组合,显著提升任务完成效率和鲁棒性。

Comments Best Paper Award Honorable Mention at NeurIPS 2025 Workshop on Bridging Language, Agent, and World Models for Reasoning and Planning

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08374 2026-03-10 cs.CV 74%

This Looks Distinctly Like That: Grounding Interpretable Recognition in Stiefel Geometry against Neural Collapse

这看起来明显像那:在Stiefel几何中基于可解释识别对抗神经崩溃

Junhao Jia, Jiaqi Wang, Yunyou Liu, Haodong Jing, Yueyi Wu, Xian Wu, Yefeng Zheng

机构 * Medical Artificial Intelligence Lab, Westlake University, Hangzhou, China(西湖大学医学人工智能实验室) Tencent Jarvis Lab, Shenzhen, China(腾讯 Jarvis 实验室) Hangzhou Dianzi University, Hangzhou, China(杭州电子科技大学) Xi’an Jiaotong University, Xi’an, China(西安交通大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

AI总结 本文提出AMP框架,利用Stiefel几何上的Riemannian优化提升原型网络的可解释性,通过正交基和空间正则化器减少原型崩溃问题,实现更准确的分类和更高的因果可信度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07758 2026-03-10 cs.CV 74%

AR2-4FV: Anchored Referring and Re-identification for Long-Term Grounding in Fixed-View Videos

AR2-4FV: 为固定视角视频中的长期接地引用而设计的锚点引用与重识别

Teng Yan, Yihan Liu, Jiongxu Chen, Teng Wang, Jiaqi Li, Bingzhuo Zhong

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

AI总结 AR2-4FV通过锚点库和重入先验提升固定视角视频中长期引用的准确率和效率。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05732 2026-03-09 cs.CV 74%

From Phase Grounding to Intelligent Surgical Narratives

从相位接地到智能手术叙述

Ethan Peterson, Huixin Zhan

机构 * New Mexico Institute of Mining and Technology(新墨西哥矿业技术研究所)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

AI总结 本文提出基于CLIP的多模态框架,通过自动创建手术时间线和叙述,减少手动标注的需要。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00732 2026-03-03 cs.RO cs.CV 74%

UniHM: Unified Dexterous Hand Manipulation with Vision Language Model

UniHM: 一体化的视觉语言模型用于统一的灵巧手操作

Zhenhao Zhang, Jiaxin Liu, Ye Shi, Jingya Wang

机构 * ShanghaiTech University(上海科技大学) InstAdapt

专题命中 视觉定位与Grounding :vision language model(title);分类 cs.CV

AI总结 UniHM通过统一的视觉语言模型实现灵巧手操作,利用开放词汇指令提升泛化能力和物理可行性。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11244 2026-02-13 cs.CV 74%

Stress Tests REVEAL Fragile Temporal and Visual Grounding in Video-Language Models

压力测试揭示视频-语言模型中脆弱的时间与视觉基础

Sethuraman T, Savya Khosla, Aditi Tiwari, Vidya Ganesh, Rakshana Jayaprakash, Aditya Jain, Vignesh Srinivasakumar, Onkar Kishor Susladkar, Srinidhi Sunkara, Aditya Shanmugham, Rakesh Vaideeswaran, Abbaas Alif Mohamed Nishar, Simon Jenni, Derek Hoiem

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Adobe Research(Adobe研究) Qualtrics iManage NVIDIA Amazon Science(亚马逊科学) Capital One

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

AI总结 REVEAL{}通过五个压力测试揭示视频-语言模型在时间与视觉基础方面的脆弱性,展示其在处理视频内容、时间序列和运动方面的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21769 2026-02-12 cs.CV 74%

H2OFlow: Grounding Human-Object Affordances with 3D Generative Models and Dense Diffused Flows

H2OFlow: 通过3D生成模型和密集扩散流接地人类-物体 affordances

Harry Zhang, Luca Carlone

机构 * MIT(麻省理工学院)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

AI总结 H2OFlow通过3D生成模型和密集扩散流学习人类-物体交互的3D affordances,无需人工标注,有效泛化至现实物体。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16038 2026-01-23 cs.AI 74%

Grounding Large Language Models in Reaction Knowledge Graphs for Synthesis Retrieval

将反应知识图谱接地于大语言模型以实现合成检索

Olga Bunkova, Lorenzo Di Fruscia, Sophia Rupprecht, Artur M. Schweidtmann, Marcel J. T. Reinders, Jana M. Weber

机构 * Department of Intelligent Systems, Delft University of Technology(智能系统系,代尔夫特理工大学) Department of Chemical Engineering, Delft University of Technology(化学工程系,代尔夫特理工大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

AI总结 本研究通过将反应路径检索转化为图查询生成问题,利用对齐示例的单样本提示提升大语言模型在合成规划中的检索准确性。

Comments Accepted at ML4Molecules 2025 (ELLIS UnConference workshop), Copenhagen, Denmark, December 2, 2025. Workshop page: https://moleculediscovery.github.io/workshop2025/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15533 2026-01-23 cs.AI 74%

From Generative Engines to Actionable Simulators: The Imperative of Physical Grounding in World Models

从生成引擎到可操作模拟器:世界模型中物理基础的必要性

Zhikang Chen, Tingting Zhu

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

AI总结 本文提出将世界模型重新定义为可操作模拟器,强调因果结构和约束意识,以提升医疗决策中的反事实推理和长期预见能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14052 2026-01-21 cs.CV 74%

Vision Also You Need: Navigating Out-of-Distribution Detection with Multimodal Large Language Model

视觉也需要:利用多模态大语言模型进行分布外检测导航

Haoran Xu, Yanlin Liu, Zizhao Tong, Jiaze Li, Kexue Fu, Yuyang Zhang, Longxiang Gao, Shuaiguang Li, Xingyu Li, Yanran Xu, Changwei Wang

机构 * Zhejiang University(浙江大学) Tsinghua University(清华大学) University of Chinese Academy of Sciences(中国科学院大学) Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences)(教育部计算电力网络与信息安全重点实验室,山东计算机科学中心(国家超算中心济南中心),齐鲁工业大学(山东科学院)) Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science(山东省计算电力互联网与服务计算重点实验室,山东省计算机科学基础研究中心) University of Electronic Science and Technology of China(电子科技大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) RWTH Aachen University(亚琛工业大学)

专题命中 视觉定位与Grounding :multimodal large language model(title);分类 cs.CV

AI总结 本文提出MM-OOD方法,利用多模态大语言模型的推理能力,通过多轮对话增强分布外检测,提升近远OOD任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06204 2026-01-19 cs.CV cs.MA 74%

Cascading multi-agent anomaly detection in surveillance systems via vision-language models and embedding-based classification

通过视觉-语言模型和基于嵌入的分类实现监视系统中的级联多代理异常检测

Tayyab Rehman, Giovanni De Gasperis, Aly Shmahell

机构 * University of L’Aquila(拉奎拉大学) SPEE S.R.L(SPEE公司)

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

AI总结 本文提出一种级联多代理框架,结合视觉-语言模型和嵌入分类,实现高效且可解释的异常检测,减少延迟并提升监控系统性能。

Comments Author email changed, Acknowlegement changes

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07795 2026-01-13 cs.CV 74%

Vision-Language Model for Accurate Crater Detection

面向准确陨石坑检测的视觉-语言模型

Patrick Bauer, Marius Schwinning, Florian Renk, Andreas Weinmann, Hichem Snoussi

机构 * University of Technology of Troyes(图卢兹技术大学) Hochschule Darmstadt(达姆施塔特应用科学大学) GMV for European Space Agency(欧洲航天局GMV) European Space Agency(欧洲航天局) Technische Hochschule Würzburg-Schweinfurt(维尔茨堡-施韦因富特技术大学)

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

AI总结 本文提出基于Vision Transformer的深度学习模型,用于在复杂月球成像条件下实现高精度陨石坑检测,通过低秩适应策略和组合损失函数提升检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02029 2026-01-06 cs.CV 74%

Leveraging 2D-VLM for Label-Free 3D Segmentation in Large-Scale Outdoor Scene Understanding

利用2D-VLM实现无标注3D分割以支持大规模户外场景理解

Toshihiko Nishimura, Hirofumi Abe, Kazuhiko Murasaki, Taiga Yoshida, Ryuichi Tanida

机构 * NTT Corporation(日本电报电话公司)

专题命中 视觉定位与Grounding :VLM(title);分类 cs.CV

AI总结 本文提出基于2D-VLM的无标注3D分割方法,通过虚拟相机投影和自然语言提示实现大规模户外场景的语义分割,支持开放词汇识别。

Comments 19

Journal ref 19th International Conference on Machine Vision Applications (MVA2025), IEICE Transactions on Information and Systems letter

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11792 2025-12-23 cs.AI 74%

Solver-Informed RL: Grounding Large Language Models for Authentic Optimization Modeling

Solver-Informed RL: 为真实优化建模奠定大语言模型基础

Yitian Chen, Jingfan Xia, Siyu Shao, Dongdong Ge, Yinyu Ye

机构 * Cardinal Operations, China(中国卡迪纳尔运营公司) Shanghai University of Finance and Economics(上海财经大学) The University of Hong Kong(香港大学) Antai School of Economics and Management, Shanghai Jiao Tong University(上海交通大学安泰经济管理学院) Department of Management Science and Engineering, Stanford University(斯坦福大学管理科学与工程系)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

AI总结 SIRL通过强化学习与外部优化求解器结合,提升大语言模型在优化建模中的准确性与实用性。

Journal ref 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02967 2025-12-16 cs.CL cs.AI cs.IR 74%

Grounding Large Language Models in Clinical Evidence: A Retrieval-Augmented Generation System for Querying UK NICE Clinical Guidelines

将大语言模型 grounded 在临床证据中:一种用于查询英国 NICE 临床指南的检索增强生成系统

Matthew Lewis, Samuel Thio, Amy Roberts, Catherine Siju, Whoasif Mukit, Rebecca Kuruvilla, Zhangshu Joshua Jiang, Niko Möller-Grell, Aditya Borakati, Richard JB Dobson, Spiros Denaxas

机构 * Institute of Health Informatics, University College London(健康信息学研究所,伦敦大学学院) Department of Biostatistics and Health Informatics, King’s College London(生物统计学与健康信息学系,伦敦国王学院) EPSRC DRIVE-Health CDT, London, U.K.(EPSRC DRIVE-Health CDT,伦敦,英国) Imperial College Healthcare NHS Trust(帝国学院医疗 NHS 信托) Black Country Healthcare NHS Foundation Trust(黑斯廷斯地区医疗 NHS 基础信托) Feldon Lane Surgery, The Dudley Group NHS Foundation Trust(费尔顿车道诊所,德比集团 NHS 基础信托) The Cleveland Clinic, London, U.K.(克利夫兰诊所,伦敦,英国) CogStack Limited(CogStack 有限公司)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

AI总结 本文提出一种基于RAG的系统,用于查询英国NICE临床指南,通过检索增强生成技术提升回答的准确性和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10195 2025-12-12 cs.CL cs.LG cs.MA 74%

AutoMedic: An Automated Evaluation Framework for Clinical Conversational Agents with Medical Dataset Grounding

AutoMedic: 一种用于具有医学数据集支撑的临床对话代理的自动化评估框架

Gyutaek Oh, Sangjoon Park, Byung-Hoon Kim

机构 * Yonsei University College of Medicine(延世大学医学院) Yonsei Institute for Digital Health(延世大学数字健康研究所) Yonsei University(延世大学) Institute of Behavioral Sciences in Medicine(医学行为科学研究所)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.LG

AI总结 AutoMedic是一种基于医学数据集的自动化评估框架,用于评估临床对话代理的多方面性能,包括准确性、效率、同理心和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06335 2025-12-10 cs.CV 74%

Harnessing Object Grounding for Time-Sensitive Video Understanding

利用物体接地提升时间敏感视频理解

Tz-Ying Wu, Sharath Nittur Sridhar, Subarna Tripathi

机构 * Intel(英特尔公司)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

AI总结 本文提出 GO-Tokenizer 以提升视频大型语言模型的时间敏感视频理解能力,通过实时编码紧凑的物体信息,提高模型性能并减少噪声影响。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22256 2025-12-01 cs.CV 74%

UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation

UMind-VL:一种通用的超声视觉-语言模型,用于统一的 grounded perception 和全面的 interpretation

Dengbo Chen, Ziwei Zhao, Kexin Zhang, Shishuang Zhao, Junjie Hou, Yaqian Wang, Nianxi Liao, Anlan Sun, Fei Gao, Jia Ding, Yuhang Liu, Dong Wang

机构 * Yizhun Medical AI Team(义诊医疗AI团队)

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

AI总结 UMind-VL 是一种通用超声视觉-语言模型,通过统一的 grounded perception 和 comprehensive interpretation 实现对医学影像的高效理解和诊断。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23258 2025-10-07 cs.CV 74%

OracleGS: Grounding Generative Priors for Sparse-View Gaussian Splatting

Atakan Topaloglu, Kunyi Li, Michael Niemeyer, Nassir Navab, A. Murat Tekalp, Federico Tombari

机构 * ETH Zürich(苏黎世联邦理工学院) Koç University(科卡大学) KUIS AI Center(KUIS人工智能中心) Technical University of Munich(慕尼黑技术大学) Google(谷歌) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments Project page available at: https://atakan-topaloglu.github.io/oraclegs/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02592 2025-10-06 cs.AI 74%

Multimodal Large Language Model Framework for Safe and Interpretable Grid-Integrated EVs

Jean Douglas Carvalho, Hugo Kenji, Ahmad Mohammad Saber, Glaucia Melo, Max Mauro Dias Santos, Deepa Kundur

专题命中 视觉定位与Grounding :multimodal large language model(title);分类 cs.AI

Comments This paper has been presented at the 2025 IEEE PES Conference on Innovative Smart Grid Technologies (ISGT 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏