arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

2026-02-11 至 2026-02-11 共收录 67 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 具身导航 16 篇

2602.09949 2026-02-11 cs.CV cs.AI 62%

Bladder Vessel Segmentation using a Hybrid Attention-Convolution Framework

利用混合注意力-卷积框架进行膀胱血管分割

Franziska Krauß, Matthias Ege, Zoltan Lovasz, Albrecht Bartz-Schmidt, Igor Tsaur, Oliver Sawodny, Carina Veil

机构 * Institute for System Dynamics in the University of Stuttgart(斯图加特大学系统动力学研究所) University Hospital Tübingen(图宾根大学医院) Department of Mechanical Engineering, Stanford University(斯坦福大学机械工程系)

专题命中 具身导航 :navigation(abstract);分类 cs.AI、cs.CV

AI总结 本文提出混合注意力-卷积框架,通过结合Transformer和CNN实现高精度膀胱血管分割,解决内窥镜数据中复杂变形和黏膜褶皱等挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13666 2026-02-11 cs.RO cs.AI 62%

DREAM: Domain-aware Reasoning for Efficient Autonomous Underwater Monitoring

DREAM:面向高效自主水下监测的领域感知方法

Zhenqi Wu, Abhinav Modi, Angelos Mavrogiannis, Kaustubh Joshi, Nikhil Chopra, Yiannis Aloimonos, Nare Karapetyan, Ioannis Rekleitis, Xiaomin Lin

机构 * Electrical Engineering, University of South Florida(佛罗里达州立大学电气工程系) Maryland Robotics Center (MRC), University of Maryland(马里兰大学机器人中心) Woods Hole Oceanographic Institution (WHOI)(伍兹霍尔海洋研究所) Mechanical Engineering, University of Delaware(德雷克塞尔大学机械工程系)

专题命中 具身导航 :robotic(abstract);分类 cs.RO、cs.AI

AI总结 DREAM通过视觉语言模型引导的自主框架,实现了高效、低耗的水下长期监测,显著提升了目标物体探测效率与覆盖范围。

Comments In Proceeding of ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10115 2026-02-11 cs.CV 57%

Quantum Multiple Rotation Averaging

量子多旋转平均

Shuteng Wang, Natacha Kuete Meli, Michael Möller, Vladislav Golyanik

机构 * Max Planck Institute for Informatics, SIC(马克斯·普朗克信息研究所,SIC) University of Siegen(施普伦次大学)

专题命中 具身导航 :robotics(abstract);分类 cs.CV

AI总结 本文提出IQARS算法,通过量子退火器解决多旋转平均问题,实现更高的旋转同步精度。

Comments 16 pages, 13 figures, 4 tables; project page: https://4dqv.mpi-inf.mpg.de/QMRA/

Journal ref International Conference on 3D Vision (3DV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09714 2026-02-11 cs.RO 57%

Fast Motion Planning for Non-Holonomic Mobile Robots via a Rectangular Corridor Representation of Structured Environments

通过结构化环境的矩形走廊表示实现非完整移动机器人的快速运动规划

Alejandro Gonzalez-Garcia, Sebastiaan Wyns, Sonia De Santis, Jan Swevers, Wilm Decré

专题命中 具身导航 :navigation(abstract);分类 cs.RO

AI总结 本研究提出了一种基于矩形走廊表示的高效运动规划框架,用于非完整移动机器人在复杂结构化环境中的快速导航。

Comments ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09076 2026-02-11 cs.RO 57%

Legs Over Arms: On the Predictive Value of Lower-Body Pose for Human Trajectory Prediction from Egocentric Robot Perception

下肢优于上肢:基于眼动机器人感知的人体轨迹预测的预测价值

Nhat Le, Daeun Song, Xuesu Xiao

机构 * Department of Computer Science, George Mason University(计算机科学系,乔治·马歇尔大学)

专题命中 具身导航 :navigation(abstract);分类 cs.RO

AI总结 该研究通过分析下肢骨骼关键点和生物力学提示,发现其在社交机器人导航中能更准确预测人体轨迹,提升导航效率。

Comments Accepted to IEEE ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08625 2026-02-11 cs.CV 57%

OpenMonoGS-SLAM: Monocular Gaussian Splatting SLAM with Open-set Semantics

OpenMonoGS-SLAM: 单目高斯点云SLAM与开放语义结合

Jisang Yoo, Gyeongjin Kang, Hyun-kyu Ko, Hyeonwoo Yu, Eunbyung Park

机构 * Sungkyunkwan University(成均馆大学) Yonsei University(延世大学)

专题命中 具身导航 :robotics(abstract);分类 cs.CV

AI总结 OpenMonoGS-SLAM通过结合3DGS与开放集语义,实现无需深度输入的单目SLAM,提升开放世界环境下的感知与建图性能。

Comments Work in progress. Project page: https://jisang1528.github.io/OpenMonoGS-SLAM/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19571 2026-02-11 cs.RO 57%

xFLIE: Leveraging Actionable Hierarchical Scene Representations for Autonomous Semantic-Aware Inspection Missions

xFLIE:利用可操作的层次场景表示进行自主语义感知检查任务

Vignesh Kottayam Viswanathan, Mario A. V. Saucedo, Sumeet Gajanan Satpute, Christoforos Kanellakis, George Nikolakopoulos

机构 * Luleå University of Technology(卢莱大学)

专题命中 具身导航 :navigation(abstract);分类 cs.RO

AI总结 xFLIE通过3DLSG实现自主语义感知检查任务的高效路径规划和导航

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09955 2026-02-11 eess.SP 50%

Doppler Effect: Analyses and Applications in Wireless Sensing and Communications

多普勒效应:无线传感与通信中的分析与应用

Lie-Liang Yang

专题命中 具身导航 :navigation(abstract)

AI总结 本文分析了多普勒效应在无线传感和通信中的影响,探讨了多种运动轨迹和相关因素,为信号频率偏移的理论基础提供支持。

Comments This document is a chapter of my next book to be published. If you have any comments, please email: lly@ecs.soton.ac.uk, which is highly appreciated

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08608 2026-02-11 eess.SY cs.SY 50%

NLoS Localization with Single Base Station Based on Radio Map

基于无线电地图的单基站非视距定位

Jiajie Xu, Yifan Guo, Xiucheng Wang, Nan Cheng, Tingting Yang

专题命中 具身导航 :navigation(abstract)

AI总结 本文提出基于无线电地图的单基站非视距定位方法,通过整合顺序信号测量与先验无线电信息,有效缓解多径效应影响,实现亚米级定位精度。

Comments The manuscript lacks a complete description of the radio map generation process, which is foundational to the localization method. We believe it is necessary to withdraw the current version to prevent the dissemination of misleading results. A corrected version will be submitted as a replacement once the issue is resolved

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09141 2026-02-11 astro-ph.IM astro-ph.CO astro-ph.EP gr-qc 50%

NIAC project report: Solar system-scale VLBI to dramatically improve cosmological distance measurements

NIAC项目报告:太阳系尺度的VLBI将大幅提高宇宙距离测量

Matthew McQuinn, Miguel Morales, Casey McGrath, Alyssa Alvarez, Katelyn Glasby, T. Joseph W. Lazio, Kiyoshi Masui, Lyujia Pan, Jonathan Pober, Huangyu Xiao

专题命中 具身导航 :navigation(abstract)

AI总结 CPS通过太阳系外VLBI技术实现宇宙距离测量,利用FRB计时推断距离,有望提高哈勃常数精度并探索暗物质和引力波

Comments 59 pages in big-ish font; NASA/NIAC Phase I study

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24654 2026-02-11 cs.CL 50%

Evolving Interactive Diagnostic Agents in a Virtual Clinical Environment

在虚拟临床环境中进化交互诊断代理

Pengcheng Qiu, Chaoyi Wu, Junwei Liu, Qiaoyu Zheng, Yusheng Liao, Haowen Wang, Yun Yue, Qianrui Fan, Shuai Zhen, Jian Wang, Jinjie Gu, Yanfeng Wang, Ya Zhang, Weidi Xie

专题命中 具身导航 :world model(abstract)

AI总结 本研究提出DiagAgent,通过强化学习在虚拟临床环境中训练,实现多轮交互式诊断,显著提升诊断准确性和检查推荐效果。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 具身推理 4 篇

2602.09856 2026-02-11 cs.CV cs.AI cs.CL cs.HC 84%

Code2World: A GUI World Model via Renderable Code Generation

Code2World: 通过可渲染代码生成实现一个GUI世界模型

Yuhao Zheng, Li'an Zhong, Yi Wang, Rui Dai, Kaikui Liu, Xiangxiang Chu, Linyuan Lv, Philip Torr, Kevin Qinghong Lin

机构 * University of Science(科学大学) AMAP, Alibaba Group(AMAP,阿里巴巴集团) University of Oxford(牛津大学) Sun Yat-sen University(中山大学)

专题命中 具身推理 :world model(title,abstract);navigation(abstract);分类 cs.AI、cs.CV

AI总结 Code2World通过可渲染代码生成实现高保真UI预测,提升Android导航性能

Comments github: https://github.com/AMAP-ML/Code2World project page: https://amap-ml.github.io/Code2World/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03213 2026-02-11 cs.CV 79%

ConsisDrive: Identity-Preserving Driving World Models for Video Generation by Instance Mask

ConsisDrive: 用于视频生成的实例掩码驾驶世界模型

Zhuoran Yang, Yanyong Zhang

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 ConsisDrive通过实例掩码注意力和损失机制提升驾驶视频生成质量及自动驾驶性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23429 2026-02-11 cs.CV 79%

Hunyuan-GameCraft-2: Instruction-following Interactive Game World Model

Hunyuan-GameCraft-2:基于指令的交互式游戏世界模型

Junshu Tang, Jiacheng Liu, Jiaqi Li, Longhuang Wu, Haoyu Yang, Penghao Zhao, Siruis Gong, Xiang Yuan, Shuai Shao, Linfeng Zhang, Qinglin Lu

机构 * Tencent Hunyuan(腾讯 Hunyuan)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 Hunyuan-GameCraft-2通过自然语言提示等多模态交互方式,实现更灵活的生成游戏世界建模,提升交互性和因果一致性。

Comments Technical Report, Project page:https://hunyuan-gamecraft-2.github.io/, Demo:https://hunyuan.tencent.com/game/game-craft

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09638 2026-02-11 cs.CV 70%

VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model

VideoAfford: 通过多模态大语言模型实现人类-物体交互视频中的3D affordance grounding

Hanqing Wang, Mingyu Liu, Xiaoyu Chen, Chengwei MA, Yiming Zhong, Wenti Yin, Yuhao Liu, Zhiqing Cui, Jiahao Yuan, Lu Dai, Zhiyuan Ma, Hui Xiong

专题命中 具身推理 :manipulation(abstract);robotic(abstract);分类 cs.CV

AI总结 VideoAfford通过多模态大语言模型实现人类-物体交互视频中的3D affordance grounding,结合动态交互先验和空间感知损失函数,提升机器人操作的可操作区域识别能力。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 机器人基础模型 1 篇

2602.09722 2026-02-11 cs.RO 70%

Rethinking Visual-Language-Action Model Scaling: Alignment, Mixture, and Regularization

重新思考视觉-语言-动作模型的扩展:对齐、混合与正则化

Ye Wang, Sipeng Zheng, Hao Luo, Wanpeng Zhang, Haoqi Yuan, Chaoyi Xu, Haiweng Xu, Yicheng Feng, Mingyang Yu, Zhiyu Kang, Zongqing Lu, Qin Jin

机构 * Renmin University of China(中国人民大学) BeingBeyond Peking University(北京大学)

专题命中 机器人基础模型 :robotics(abstract);robotic(abstract);分类 cs.RO

AI总结 本研究重新审视视觉-语言-动作模型的扩展问题,探讨了对齐、混合与正则化在机器人控制中的关键作用,挑战了传统假设并提供实践指导。

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 模仿学习与强化学习 3 篇

2602.10044 2026-02-11 cs.LG cs.AI cs.SY eess.SY 81%

Optimistic World Models: Efficient Exploration in Model-Based Deep Reinforcement Learning

乐观世界模型:基于模型的深度强化学习中的高效探索

Akshay Mete, Shahid Aamir Sheikh, Tzu-Hsiang Lin, Dileep Kalathil, P. R. Kumar

机构 * Department of Electrical Computer Engineering, Texas A\&M University, College Station, Texas, USA

专题命中 模仿学习与强化学习 :world model(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出乐观世界模型,通过引入乐观动态损失提升深度强化学习中的高效探索能力,显著提高样本效率和累积回报。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13761 2026-02-11 cs.AI cs.CL 57%

THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning

THOR:通过强化学习进行工具集成的分层优化以实现数学推理

Qikai Chang, Zhenrong Zhang, Pengfei Hu, Jun Du, Jiefeng Ma, Yicheng Pan, Jianshu Zhang, Quan Liu, Jianqing Gao

机构 * University of Science and Technology of China(中国科学技术大学) iFLYTEK Research(iFLYTEK研究院)

专题命中 模仿学习与强化学习 :manipulation(abstract);分类 cs.AI

AI总结 THOR通过强化学习实现工具集成的分层优化,提升数学推理和代码生成能力。

Comments 22 pages, 13 figures, ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10357 2026-02-11 cs.AI 57%

Optimus-3: Dual-Router Aligned Mixture-of-Experts Agent with Dual-Granularity Reasoning-Aware Policy Optimization

Optimus-3: 具有双粒度推理感知策略优化的双路由对齐专家混合代理

Zaijing Li, Yuquan Xie, Rui Shao, Gongwei Chen, Weili Guan, Dongmei Jiang, Yaowei Wang, Liqiang Nie

机构 * School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen Campus)(计算机科学与技术学院,哈尔滨工业大学(深圳校区)) Pengcheng Laboratory(鹏城实验室) School of Information Science and Technology, Harbin Institute of Technology (Shenzhen Campus)(信息科学与技术学院,哈尔滨工业大学(深圳校区)) Shenzhen Loop Area Institute(深圳河套学院)

专题命中 模仿学习与强化学习 :embodied AI(abstract);分类 cs.AI

AI总结 Optimus-3通过双路由对齐专家混合架构和双粒度推理感知策略优化,实现了系统1与系统2的协同,提升了开放性任务的解决能力。

Comments 16 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 人机交互与遥操作 1 篇

2602.09287 2026-02-11 cs.RO cs.HC 57%

Disambiguating Anthropomorphism and Anthropomimesis in Human-Robot Interaction

在人机交互中区分拟人化与拟人模仿

Minja Axelsson, Henry Shevlin

机构 * University of Cambridge(剑桥大学)

专题命中 人机交互与遥操作 :robotics(abstract);分类 cs.RO

AI总结 本文区分了人机交互中拟人化与拟人模仿的概念,明确两者责任方,为未来研究提供理论基础。

Comments 6 pages, 0 figures. Accepted at Human-Robot Interaction (HRI) conference (ACM/IEEE) 2026, Late-Breaking Reports

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 机器人数据与评测 15 篇

2602.10093 2026-02-11 cs.RO 83%

UniVTAC: A Unified Simulation Platform for Visuo-Tactile Manipulation Data Generation, Learning, and Benchmarking

UniVTAC:一种用于视觉-触觉操作数据生成、学习和评估的统一仿真平台

Baijun Chen, Weijie Wan, Tianxing Chen, Xianda Guo, Congsheng Xu, Yuanyang Qi, Haojie Zhang, Longyan Wu, Tianling Xu, Zixuan Li, Yizhe Wu, Rui Li, Xiaokang Yang, Ping Luo, Wei Sui, Yao Mu

机构 * ScaleLab, Shanghai Jiao Tong University(上海交通大学ScaleLab) D-Robotics ViTai Robotics The University of Hong Kong(香港大学) Nanjing University(南京大学) Shenzhen University(深圳大学) Wuhan University(武汉大学) Fudan University(复旦大学) Tsinghua University(清华大学)

专题命中 机器人数据与评测 :manipulation(title,abstract);robotic(abstract);分类 cs.RO

AI总结 UniVTAC提出了一种统一的视觉-触觉数据生成平台,通过训练触觉中心的编码器和基准任务提升机器人操作的成功率。

Comments Website: https://univtac.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09367 2026-02-11 cs.RO 83%

CAPER: Constrained and Procedural Reasoning for Robotic Scientific Experiments

CAPER:用于机器人科学实验的约束和程序性推理

Jinghan Yang, Jingyi Hou, Xinbo Yu, Wei He, Yifan Wu

机构 * University of Science and Technology Beijing(北京科技大学) Beijing Information Science and Technology University(北京信息科技大学)

专题命中 机器人数据与评测 :robotic(title,abstract);manipulation(abstract);分类 cs.RO

AI总结 CAPER通过约束和程序性推理框架提升机器人在科学实验中的可控性、鲁棒性和数据效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21797 2026-02-11 cs.CV 77%

MoWM: Mixture-of-World-Models for Embodied Planning via Latent-to-Pixel Feature Modulation

MoWM: 通过潜在到像素特征调制的混合世界模型实现具身规划

Yangcheng Yu, Xin Jin, Yu Shang, Xin Zhang, Haisheng Su, Wei Wu, Yong Li

机构 * Tsinghua University(清华大学) Manifold AI Shanghai Jiao Tong University(上海交通大学)

专题命中 机器人数据与评测 :robotics(abstract);manipulation(abstract);world model(abstract);分类 cs.CV

AI总结 MoWM通过融合潜在世界模型与像素特征,提升具身规划中动作解码的精度与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09772 2026-02-11 cs.RO 74%

Design and Evaluation of an Assisted Programming Interface for Behavior Trees in Robotics

行为树在机器人中的辅助编程接口设计与评估

Jonathan Styrud, Matteo Iovino, Rebecca Stower, Mart Kartašev, Mikael Norrlöf, Mårten Björkman, Christian Smith

专题命中 机器人数据与评测 :robotics(title);分类 cs.RO

AI总结 本文提出BETR-GUI,结合AI助手与拖放编辑器,通过整合多种技术提升机器人行为树编程效率,实验证明人类用户优于纯AI助手。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18083 2026-02-11 cs.LG cs.RO 73%

What Do You Need for Compositional Generalization in Diffusion Planning?

在扩散规划中你需要什么才能实现组合泛化?

Quentin Clark, Florian Shkurti

机构 * Department of Computer Science, University of Toronto(计算机科学系,多伦多大学)

专题命中 机器人数据与评测 :manipulation(abstract);navigation(abstract);分类 cs.RO、cs.LG

AI总结 本文提出Eq-Net架构,通过位移等变性、局部感受野和推断选择实现扩散规划的组合泛化,展示其在导航和操作任务中的有效性。

Comments 8 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10102 2026-02-11 cs.CV 70%

VideoWorld 2: Learning Transferable Knowledge from Real-world Videos

VideoWorld 2: 从真实世界视频中学习可迁移的知识

Zhongwei Ren, Yunchao Wei, Xiao Yu, Guixun Luo, Yao Zhao, Bingyi Kang, Jiashi Feng, Xiaojie Jin

机构 * Beijing Jiaotong University(北京交通大学)

专题命中 机器人数据与评测 :robotics(abstract);manipulation(abstract);分类 cs.CV

AI总结 VideoWorld 2通过动态增强的潜在动态模型从真实世界视频中学习可迁移知识,显著提升了任务成功率和长周期推理能力。

Comments Code and models are released at: https://maverickren.github.io/VideoWorld2.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09893 2026-02-11 cs.RO cs.AI 62%

TaCo: A Benchmark for Lossless and Lossy Codecs of Heterogeneous Tactile Data

TaCo: 一种用于异构触觉数据无损和有损编解码器的基准测试

Zhengxue Cheng, Yan Zhao, Keyu Wang, Hengdi Zhang, Li Song

机构 * Shanghai Jiao Tong University(上海交通大学) Paxini Tech.(帕辛尼科技)

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.RO、cs.AI

AI总结 TaCo提出了一种用于评估触觉数据编解码器性能的基准测试,通过评估30种压缩方法,包括神经编解码器,验证了无损和有损压缩方案在多个任务中的优越性能。

Comments 27 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19713 2026-02-11 cs.CV cs.RO 62%

VIMD: Monocular Visual-Inertial Motion and Depth Estimation

VIMD:单目视觉-惯性运动和深度估计

Saimouli Katragadda, Guoquan Huang

机构 * University of Delaware (UD) College of Engineering(德克萨斯大学达勒姆分校工程学院) University of Delaware(德克萨斯大学达勒姆分校) Robot Perception and Navigation Group (RPNG)(机器人感知与导航小组)

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.RO、cs.CV

AI总结 VIMD通过多视图信息迭代优化像素尺度,实现高效且鲁棒的单目视觉-惯性深度估计,适用于资源受限的3D视觉感知场景。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22516 2026-02-11 cs.LG cs.CV 62%

Ice-FMBench: A Foundation Model Benchmark for Sea Ice Type Segmentation

Ice-FMBench: 一个用于海冰类型分割的基础模型基准

Samira Alkaee Taleghan, Morteza Karimzadeh, Andrew P. Barrett, Walter N. Meier, Farnoush Banaei-Kashani

机构 * University of Colorado Denver(科罗拉多大学丹佛分校) University of Colorado Boulder(科罗拉多大学波德分校) National Snow and Ice Data Center (NSIDC), CIRES, University of Colorado Boulder(国家冰雪数据研究中心(NSIDC)、CIRES、科罗拉多大学波德分校)

专题命中 机器人数据与评测 :navigation(abstract);分类 cs.CV、cs.LG

AI总结 Ice-FMBench为海冰类型分割任务提供了一个基准框架,评估基础模型的性能,并通过多教师知识蒸馏方法提升模型的时空可转移性。

Journal ref ACM ACM SIGSPATIAL PoIDS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21872 2026-02-11 eess.IV cs.LG 57%

Targeted Unlearning Using Perturbed Sign Gradient Methods With Applications On Medical Images

利用扰动符号梯度方法的目标遗忘与医疗图像应用

George R. Nahass, Zhu Wang, Homa Rashidisabet, Won Hwa Kim, Sasha Hubschman, Jeffrey C. Peterson, Chad A. Purnell, Pete Setabutr, Ann Q. Tran, Darvin Yi, Sathya N. Ravi

机构 * Department of Biomedical Engineering(生物医学工程系) Department of Ophthalmology(眼科学系) University of Illinois Chicago(伊利诺伊大学芝加哥分校) Department of Computer Science(计算机科学系) Computer Science and Engineering(计算机科学与工程) Pohang University of Science and Technology, South Korea(韩国釜山科学技术大学) Department of Plastic and Reconstructive Surgery(整形外科与重建外科系)

专题命中 机器人数据与评测 :world model(abstract);分类 cs.LG

AI总结 本文提出了一种基于扰动符号梯度的方法,用于医疗图像中的目标遗忘,通过可调损失设计和模型组合策略,在遗忘与保留之间取得平衡,优于现有基线方法。

Comments 39 pages, 12 figures, 11 tables, 3 algorithms

Journal ref Transactions on Machine Learning Research 2025, https://openreview.net/forum?id=XE0bJg6sQN

详情

展开后加载摘要…

URL PDF HTML 收藏