arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7473 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7473 篇

2510.17651 2025-10-21 cs.CV cs.AI cs.LG 80%

Frugal Federated Learning for Violence Detection: A Comparison of LoRA-Tuned VLMs and Personalized CNNs

Sébastien Thuau, Siba Haidar, Ayush Bajracharya, Rachid Chelouah

机构 * esieaLab(esiea实验室) ESIEA(ESIEA学院) ETIS Laboratory(ETIS实验室) CNRS(法国国家科学研究中心) UMR8051(UMR8051研究中心) University of CY Cergy(CY塞克大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);LLaVA(abstract);分类 cs.CV、cs.AI、cs.LG

Comments 7 pages, 1 figure, FLTA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11616 2025-08-18 cs.CV cs.AI cs.CL cs.LG 80%

Controlling Multimodal LLMs via Reward-guided Decoding

Oscar Mañas, Pierluca D'Oro, Koustuv Sinha, Adriana Romero-Soriano, Michal Drozdzal, Aishwarya Agrawal

机构 * Mila - Quebec AI Institute(魁北克AI研究院) Université de Montréal(蒙特利尔大学) McGill University(麦吉尔大学) Meta FAIR Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Published at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23573 2025-04-01 cs.CV cs.AI cs.LG 80%

DASH: Detection and Assessment of Systematic Hallucinations of VLMs

Maximilian Augustin, Yannic Neuhaus, Matthias Hein

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);LLaVA(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03151 2025-01-07 cs.AI cs.CV cs.LG 80%

Large language models for artificial general intelligence (AGI): A survey of foundational principles and approaches

Alhassan Mumuni, Fuseini Mumuni

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23143 2026-08-25 cs.CV 新提交 79%

An end-to-end-trained vision-language model for native-language prostate pathology report generation

用于生成本土语言前列腺病理报告的端到端训练视觉-语言模型

Christian Grashei, Fabian Gülhan, Maximilian Legnar, Fabian Stögbauer, Cleo-Aron Weis, Carolin Mogler, Peter Schüffler

机构 * Technical University of Munich(慕尼黑工业大学) Munich Data Science Institute(慕尼黑数据科学研究所) Munich Center for Machine Learning(慕尼黑机器学习中心) University Hospital Heidelberg(海德堡大学医院) Heidelberg University(海德堡大学) Interdisciplinary Center for Scientific Computing (IWR)(跨学科科学计算中心(IWR))

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

AI总结 该研究提出语言独立的切片级视觉-语言框架,用自动化流水线生成17344对图像-文本对,实现德语前列腺病理报告生成,恶性肿瘤检测F1达96.2%,可支持机构用自有档案训练本土语言报告模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21878 2026-08-25 cs.CV 新提交 79%

ViSMoE: Visual-Aware Sparse Mixture-of-Experts for Embodied Referring Expression Grounding

ViSMoE:面向具身指代表达接地的视觉感知稀疏混合专家模型

Shuo Feng, Piji Li

机构 * College of Artificial Intelligence(人工智能学院) Nanjing University of Aeronautics and Astronautics(南京航空航天大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 针对现有具身指代表达接地方法无法区分视角与物体导致表示模糊的问题,提出带视觉感知路由策略的ViSMoE框架,在REVERIE和SOON数据集上性能优于现有SOTA方法。

Comments Accepted by ICANN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15382 2026-08-25 cs.AI cs.CL cs.IR q-bio.QM 版本更新 79%

Framework for Grounding Healthcare LLMs in a Causal Knowledge Graph: A Cardiovascular Example Pilot

将医疗大语言模型(LLM)基于因果知识图谱:框架、指标与心血管试点研究

Ummara Mumtaz, Aimen Noor, Awais Ahmed

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 本研究提出以因果知识图谱为核心的医疗LLM评估框架,在心血管试点中验证其有效性,发现集成条件C4在因果推理相关指标上表现最优,未基于图的C1原始干预准确性最高但缺乏因果与证据基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20392 2026-08-24 cs.CL cs.AI 新提交 79%

Evaluation-as-Search: Adaptive Discovery of Grounding Failures in Meeting Assistants

评估即搜索:会议助手接地故障的自适应发现

Sami Khairy, Yasaman Hosseinkashi, Vishak Gopal, Ross Cutler

机构 * Microsoft(微软公司)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 本文提出EaS方法构建MeetingProbe基准,发现会议助手故障多源于话语语用挑战,其自适应搜索比随机探测的故障发现率高2.5倍,且公开了该基准以支持可复现评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20336 2026-08-21 cs.CV 新提交 79%

WithEveryone: Unified Planning and Identity Grounding for Group Image Generation

WithEveryone:面向群体图像生成的统一规划与身份定位

Hengyuan Xu, Qixun Wang, Yiji Cheng, Miles Yang, Zhao Zhong, Wei Cheng, Xingjun Ma, Yu-gang Jiang

机构 * Fudan University(复旦大学) Hunyuan, Tencent(腾讯混元) The University of Hong Kong(香港大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 针对多人物群体图像生成的身份保留难题,提出WithEveryone框架,通过布局定位身份损失等技术,提升了人脸相似度并降低伪影,实现高身份覆盖率与低重复率。

Comments Project Page: doby-xu.github.io/WithEveryone/ ;Code will be released: github.com/Doby-Xu/WithEveryone/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19553 2026-08-21 cs.CV 新提交 79%

Where Grounding Accuracy Lives on the IoU Curve: Label-Free Inference-Time Boundary Refinement

IoU曲线上定位精度的核心所在:无标注推理时的边界优化

Bo Ma

机构 * Auckland University of Technology(奥克兰理工大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 本文提出无标注推理时的边界优化(LFPR)方法,在多个视觉-语言定位数据集上提升了IoU曲线各精度指标,证明指代选择与边界精度部分可分离,为定位精度优化提供了新思路。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07932 2026-08-21 cs.CV 版本更新 79%

SportsGrounder: Proposal-Aided Interleaved Grounding for Dense Sports Video Reasoning

SportsGrounder:用于密集体育视频推理的提案辅助交错定位框架

Yizhi Li, Jiawei Jiang, Guanhong Wang, Yingcai Wu, Gaoang Wang

机构 * Zhejiang University(浙江大学) Beijing Jiaotong University(北京交通大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 针对LMMs难以处理密集体育视频细粒度推理的问题,提出SportsGrounder框架,结合IGF机制与AAS模块,在SoccerNet等数据集上实现最优准确率。

Comments ACMMM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19011 2026-08-20 cs.CR cs.AI 新提交 79%

From Threat Intelligence to Detection: Knowledge-driven Enrichment and Template-based Rule Grounding for Automated Sigma Rule Generation

从威胁情报到检测:知识驱动的丰富化与基于模板的规则落地用于自动化Sigma规则生成

Sepehr Ghaffarzadegan, Boubakr Nour, Makan Pourzandi, Mourad Debbabi, Chadi Assi

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 本研究设计了AUTOSIGMA,结合知识驱动丰富化等技术,将非结构化CTI报告自动生成Sigma规则,其在多项指标上优于其他方案。

Comments Submitted for publication and currently under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12746 2026-08-20 cs.CV cs.CL 版本更新 79%

Dual-Stream Cross-Anchor Correction Grounding Long-Form Captions and the Domain Limits of Object-Level Anchors

双流跨锚校正:接地长文本描述与对象级锚点的域限制

Lingkai Bu, Qian Gao, Jun Fan, Guohui Ding, Zhenyu Yang, Yuteng Xiao, Jinyi Liang

专题命中 视觉定位与Grounding :grounding(title);multimodal large language model(abstract);分类 cs.CV

AI总结 针对多模态大语言模型的长文本描述对象幻觉问题,本文提出双流跨锚校正方法,通过耦合感知流与认知流提升精度,在长文本场景下实现最优性能,且存在域条件性限制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17328 2026-08-19 cs.CV 新提交 79%

MS-MFAD : Multimodal large language models for Face Anti-spoofing Detection

MS-MFAD:用于人脸活体检测的多模态大语言模型

Xiaoyong Yu, Rongzhen Li, Shuming Shi, Xinge You

机构 * Mashang Consumer Finance Co., Ltd.(马上消费金融股份有限公司) Huazhong University of Science and Technology(华中科技大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);分类 cs.CV

AI总结 针对人脸活体检测的复合威胁与现有方法瓶颈,提出基于多模态大语言模型MFAD的可解释系统,通过细粒度像素-语义锚定机制与少量语义标注实现优异检测性能,鲁棒性与效率更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27155 2026-08-19 cs.AI cs.CL cs.HC 版本更新 79%

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

OmegaUse-OfficeVal:基于经济基准的长周期办公套件任务LLM智能体评测基准

Jingbo Zhou, Yusai Zhao, Qi Bao, Jingjia Cao, Zhenghai Chen, Chang Gao, Kaiqi Guo, Muxin Guo, Mingxuan Li, Xinjiang Lu, Yanru Ma, Yixiong Xiao, Zenghui Zhang, Le Zhang, Hua Wu

机构 * Baidu Inc.(百度公司)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 OmegaUse-OfficeVal是带经济基准的长周期办公套件任务LLM智能体评测基准,含100项任务,评测发现前沿LLM成本低速度快但质量未达人类水平,代码与数据集已开源。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15419 2026-08-18 cs.CV 新提交 79%

ArtLang: Structured Language-to-Kinematics Grounding for Articulated 3D Actuation

ArtLang:面向铰接3D驱动的结构化语言到运动学的绑定

Sylvia Yuan, Dan Wang, Ravi Ramamoorthi, Xinrui Cui

机构 * University of California San Diego(加利福尼亚大学圣迭戈分校) University of North Texas(北得克萨斯大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 研究针对铰接物体控制的语义匿名问题,提出ArtLang框架,通过语义-运动学铰接图等技术实现开放词汇语言控制,经多类实验验证了可靠的语言绑定与连续铰接控制能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14228 2026-08-17 cs.LG 新提交 79%

AutoSchema: Live Schema Grounding for Agentic Text-to-Sparql over Heterogeneous Knowledge Graphs

AutoSchema:面向异构知识图谱的智能体文本转SPARQL的实时模式接地

Yiming Zhang, Koji Tsuda

机构 * The University of Tokyo(东京大学) National Institute for Materials Science(国立材料科学研究所) RIKEN Center for Advanced Intelligence Project(理化学研究所高级智能项目中心)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.LG

AI总结 提出无需训练的AutoSchema框架,用于异构知识图谱的智能体文本转SPARQL的实时模式接地,在多项生物医学KGQA等任务中优于TogoMCP,可支持未记录RDF图谱。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12748 2026-08-14 cs.CV 新提交 79%

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding

扩展表示多样性:用于视觉定位的调制注意力与重构正则化

Junyi Hu, Tian Bai, Fengyi Wu, Yian Huang, Wei Wen, Zaoli Li, Junli Lin, Xingchen Li, Zhenming Peng, Yi Zhang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 针对指代表达理解模型跨数据集泛化能力有限的问题,提出含mACH与JEPA辅助流的架构及Objects365-Caption数据,实现了强泛化与具竞争力的REC性能。

Comments 21 pages, 10 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12683 2026-08-14 cs.RO cs.CV 新提交 79%

FUSE: Active Functional Affordance Grounding through Adaptive Semantic-Geometric Evidence Acquisition

FUSE:通过自适应语义-几何证据获取实现主动功能可供性接地

Zhou Chen, Sathyanarayanan N. Aakur

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 该研究提出主动功能可供性接地任务,构建FUSE框架结合不确定性驱动探索与摊销规划器,基于Habitat基准验证其在非神示接地中性能最优且计算量降低1.33倍。

Comments Under review. 15 Pages. 9 tables, 3 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11839 2026-08-13 cs.LG 版本更新 79%

Grounding Large Language Models as Generalizable Policies in Network Control

大语言模型作为网络优化的通用策略

Duo Wu, Linjia Kang, Zhimin Wang, Fangxin Wang, Wei Zhang, Chongbo Sun, Xuefeng Tao, Wei Yang, Le Zhang, Wenwu Zhu, Peng Cui, Zhi Wang

机构 * Bytedance(字节跳动) Shenzhen International Graduate School(深圳国际研究生院) Tsinghua University(清华大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Department of Computer Science and Technology(计算机科学与技术系)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.LG

AI总结 本文提出Trailblazer框架,利用大语言模型实现跨任务和环境的通用网络策略,通过网络对齐和策略协作机制提升效率与泛化能力。

Comments Arxiv version. Official version has been submitted to IEEE Transactions on Mobile Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10756 2026-08-12 cs.RO cs.CV 新提交 79%

Embodied Multimodal Grounding for Open-Vocabulary Mobile Manipulation via Semantic 3D Gaussian Splatting

基于语义三维高斯溅射的开放词汇移动操作具身多模态定位

Huosen Ou, Dongni Song, Yuncong Wang, Tao Zhou, Yiding Ji

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Midea Group(美的集团) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 本文针对家庭场景开放词汇移动操作的目标定位问题,提出整合Semantic-3DGS的具身多模态框架,经50次真实机器人试验,在长 horizon、杂乱场景等任务中优于PointVLA等基线方法,提升了操作鲁棒性。

Comments 9 pages, 11 figures. Accepted to ACM Multimedia 2026 (MM '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17162 2026-08-12 cs.AI q-bio.GN 版本更新 79%

JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures

JEPA-DNA:通过联合嵌入预测架构夯实基因组基础模型

Ariel Larey, Elay Dahan, Amit Bleiweiss, Raizy Kellerman, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, Nicole Bussola, Simon Lee, Shane O'Connell, Dung Hoang, Marissa Wirth, Alexander W. Charney, Nati Daniel, Yoli Shavit

机构 * Applied AI Architecture, NVIDIA, Israel(NVIDIA应用人工智能架构,以色列) Worldwide Field Ops, NVIDIA, Israel(NVIDIA全球现场运营,以色列) Developer Programs, NVIDIA, Israel(NVIDIA开发者计划,以色列) Cancer Research Center and Wohl Institute of Translational Medicine, Sheba Medical Center, Tel Hashomer, Israel(癌症研究中心和Wohl转化医学研究所,Sheba医疗中心,Tel Hashomer,以色列) Windreich Department of AI and Human Health, Icahn School of Medicine at Mount Sinai, New York, USA(AI与人类健康风reich部门,Mount Sinai医学中心,纽约,美国)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 提出JEPA-DNA框架,将联合嵌入预测架构与生成式目标结合,通过潜在空间监督全局序列嵌入,实现从令牌恢复到语义对齐的转变,在17项基因组基准任务上提升线性探测和零样本性能,达到新最优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09143 2026-08-11 cs.CV 新提交 79%

UniMoFlow: Grounding Instruction-Driven 3D Human Motion Editing in Generation

UniMoFlow:将指令驱动的3D人体动作编辑建立在生成基础上

Yilei Hua, Beibei Jing, Ce Zheng, Hanyu Zhou, Yawei Luo, Wei Yang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 UniMoFlow将指令驱动3D人体动作编辑建立在文本到动作生成基础上,通过构建Omni-MoEdit数据集、提出UniMoFlow模型与SAFE方法,提升了编辑效果与生成质量

Comments 18 pages, including supplementary material; 8 figures and 7 tables. Code: https://github.com/Yilei-Hua/UniMoFlow. Submitted to AAAI 2027

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08929 2026-08-11 cs.CV 新提交 79%

Zero-shot 2D Grounding with Novel Affordance Types

基于新型可供性类型的零样本二维定位

Haomeng Zhang, Raymond A. Yeh

机构 * Purdue University(普渡大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 本文提出新型可供性类型的零样本二维定位任务,构建对应基准,提出AffordAnything及可训练变体AffordAnything+,在AGD20K-NAT基准上较SOTA方法OOAL的IoU@0.4提升12.3%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12602 2026-08-11 cs.CV 版本更新 79%

Decouple and Reason: Anatomically Guided Two-Stage Voxel-Level Grounding of Free-Text Findings in 3D Chest CT

解耦与推理:3D胸部CT中自由文本发现的解剖学引导两阶段体素级定位

Kwang-Hyun Uhm, Inhwa Son, Sung-Jea Ko

机构 * Department of Artificial Intelligence, Gachon University(韩国加图立大学人工智能系) MEDAI

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 研究3D胸部CT中自由文本发现的体素级定位难题,提出解耦框架,分病变分割和文本-体积推理两阶段,利用解剖学引导解决空间模糊性,在基准测试中取得领先,证明解耦是处理该复杂性的有效范式。

Comments Accepted to MICCAI 2026 (Spotlight Presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04935 2026-08-10 cs.CV 版本更新 79%

Unleashing the Potential of Vision-Language Models for Generalizable AI-Generated Image Detection

释放视觉-语言模型在通用人工智能生成图像检测中的潜力

Weihan Cai, Hao Tan, Zichang Tan, Jun Wan, Xinping Gao

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

AI总结 该研究针对AI生成图像检测,发现视觉-语言模型PE比DINOv3更具潜力,提出语义原型校准(SPC)方法得到PE-SPC,在多基准测试中达到新的最先进性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05569 2026-08-07 cs.CV 新提交 79%

CoordRefer: Coordinate-Aware 3D Visual Grounding from Multiview Images

CoordRefer:基于多视图图像的坐标感知三维视觉定位

Haijie Li, Jiaxin Zhang, Dave Zhenyu Chen, Youyu Chen, Yanmin Wu, Jian Zhang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 本文提出CoordRefer框架,解耦坐标帧选择与坐标条件三维定位,在ScanRefer数据集上实现定位精度提升,性能优于坐标不可知基线及部分显式三维输入方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13621 2026-08-07 cs.CV 版本更新 79%

Visual Intention Grounding for Egocentric Assistants

面向第一人称视角助手的视觉意图定位

Pengzhan Sun, Junbin Xiao, Tze Ho Elden Tse, Yicong Li, Arjun Akula, Angela Yao

机构 * National University of Singapore(新加坡国立大学) Google DeepMind(谷歌DeepMind)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 本文提出首个第一人称视觉意图定位数据集EgoIntention,及推理至定位(RoG)指令微调方法,解决多模态模型在第一人称视角下的意图定位问题,实现了统一的视觉定位能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03763 2026-08-05 cs.CV 新提交 79%

TDVR: Joint Text Disambiguation and Viewpoint Reasoning for Zero-Shot 3D Visual Grounding

TDVR:用于零样本3D视觉定位的联合文本消歧与视点推理框架

Qingxi Du, Junbo Wang, Yuke Li, Yining Zhu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 本文提出TDVR框架,通过联合文本消歧与视点推理解决零样本3D视觉定位中的文本歧义与视点不足问题,在ScanRefer数据集上的Acc@0.25和Acc@0.5指标较现有最优方法分别提升15.25%和14.46%

Comments 10 pages, 5 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03322 2026-08-05 cs.CV 新提交 79%

LocAnyMed: Vision-Language Grounding for Multimodal Medical Images

LocAnyMed:面向多模态医学图像的视觉-语言定位

Zihan Wang, Tong Liu, Zhiwei Wang, Tao Huang, Wentao Jiang, Sihan Ma, Shanshan Ye, Xiaohui Yang, Jing Zhang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 该研究构建多模态医学视觉定位数据集LocAnyMed-200K及理由增强子集LocAnyMed-CoT-20K,通过微调LocateAnything-3B提升医学定位性能,为相关研究提供统一基础。

Comments Technical report; work in progress. 28 pages, 5 figures, and 16 tables. Code: https://github.com/MiliLab/LocAnyMed

详情

展开后加载摘要…

URL PDF HTML 收藏