arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2816 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2816 篇

2602.15767 2026-02-18 cs.RO cs.AI cs.HC 57%

Robot-Assisted Social Dining as a White Glove Service

机器人辅助社交用餐作为白手套服务

Atharva S Kashyap, Ugne Aleksandra Morkute, Patricia Alves-Oliveira

机构 * University of Michigan(密歇根大学) Leiden University(莱顿大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 本研究通过参与设计和AI工具,提出机器人辅助社交用餐的白手套服务理念,强调多模态输入、情境敏感行为及角色扩展,以提升残疾人在野外社交用餐中的独立性和尊严。

Comments 20 pages, 9 figures. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15294 2026-02-18 cs.AI 57%

EAA: Automating materials characterization with vision language model agents

EAA: 用视觉语言模型代理自动化材料表征

Ming Du, Yanqi Luo, Srutarshi Banerjee, Michael Wojcik, Jelena Popovic, Mathew J. Cherukara

机构 * Argonne National Laboratory(阿贡国家实验室) Advanced Photon Source(先进光子源) Data Science and Learning Division(数据科学与学习 division) Department of Radiation Oncology(放射肿瘤学系) Northwestern University(西北大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 EAA利用视觉语言模型代理自动化材料表征,通过多模态推理和工具增强操作提升显微镜实验效率,减少操作负担并降低用户专业门槛。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14093 2026-02-17 cs.AI cs.LG 57%

GUI-GENESIS: Automated Synthesis of Efficient Environments with Verifiable Rewards for GUI Agent Post-Training

GUI-GENESIS: 自动合成具有可验证奖励的高效环境以实现GUI代理训练

Yuan Cao, Dezhi Ran, Mengzhou Wu, Yuzhe Guo, Xin Chen, Ang Li, Gang Cao, Gong Zhi, Hao Yu, Linyi Li, Wei Yang, Tao Xie

机构 * Key Lab of HCST (PKU), MOE SCS, Peking University, Beijing, China Tencent Inc., Shenzheng, China Hong Kong University of Science Department of Computer Science, University of Texas at Dallas, Richardson, USA School of Computing Science, Simon Fraser University, Burnaby, BC, Canada

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 GUI-GENESIS通过自动合成高效GUI训练环境并提供可验证奖励,显著提升了GUI代理的训练效率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14048 2026-02-17 cs.RO cs.CV cs.GR 57%

ProAct: A Dual-System Framework for Proactive Embodied Social Agents

ProAct:一种双系统框架用于主动具身社交代理

Zeyi Zhang, Zixi Kang, Ruijie Zhao, Yusen Feng, Biao Jiang, Libin Liu

机构 * Peking University(北京大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 ProAct通过双系统框架实现主动具身社交代理,结合低延迟行为系统与慢速认知系统,提升交互的主动性和社会参与度。

Comments Project Page: https://proactrobot.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14003 2026-02-17 cs.AI 57%

Prompt-Driven Low-Altitude Edge Intelligence: Modular Agents and Generative Reasoning

基于提示的低空边缘智能:模块化代理与生成推理

Jiahao You, Ziye Jia, Chao Dong, Qihui Wu

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

AI总结 本文提出P2AECF框架,通过模块化代理和生成推理实现灵活、高效和适应的低空边缘智能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10080 2026-02-17 cs.CV 57%

BEVTraj: Map-Free End-to-End Trajectory Prediction in Bird's-Eye View with Deformable Attention and Sparse Goal Proposals

BEVTraj: 无地图端到端鸟瞰图轨迹预测方法,采用可变形注意力和稀疏目标提案

Minsang Kong, Myeongjun Kim, Sang Gu Kang, Hejiu Lu, Yupeng Zhong, Sang Hun Lee

机构 * Department of Automobile and IT Convergence, Kookmin University(汽车与IT融合系,韩国釜山大学) Department of Automotive Engineering, Kookmin University(汽车工程系,韩国釜山大学) Graduate School of Automobile and Mobility, Kookmin University(汽车与移动研究生院,韩国釜山大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 BEVTraj通过可变形注意力和稀疏目标提案实现无地图端到端鸟瞰图轨迹预测,提升自动驾驶的鲁棒性和灵活性。

Comments Submitted to IEEE Transactions on Intelligent Transportation Systems (under review)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12083 2026-02-13 cs.AI cs.LO 57%

Differentiable Modal Logic for Multi-Agent Diagnosis, Orchestration and Communication

可微模态逻辑用于多智能体诊断、协调与通信

Antonin Sulc

机构 * Lawrence Berkeley National Lab(伯克利劳伦斯国家实验室) Berkeley(伯克利)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

AI总结 本文提出可微模态逻辑,通过神经符号方法实现多智能体系统的调试与协调,结合知识注入和多模态推理,提升智能体间的信任与因果推理能力。

Comments 29 pages, 8 figures, 8 tables, Tutorial at 3rd International Conference on Neuro-Symbolic Systems (NeuS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11750 2026-02-13 cs.SE cs.AI cs.HC 57%

AmbiBench: Benchmarking Mobile GUI Agents Beyond One-Shot Instructions in the Wild

AmbiBench:在真实场景中超越单次指令的移动GUI代理基准测试

Jiazheng Sun, Mingxuan Li, Yingying Zhang, Jiayang Niu, Yachen Wu, Ruihan Jin, Shuyu Lei, Pengrongrui Tan, Zongyu Zhang, Ruoyi Wang, Jiachen Yang, Boyu Yang, Jiacheng Liu, Xin Peng

机构 * Fudan University(复旦大学) Jilin University(吉林大学)

专题命中 多模态Agent :MLLM(abstract);分类 cs.AI

AI总结 AmbiBench是首个针对移动GUI代理在真实场景中超越单次指令的双向意图对齐的基准测试,通过引入清晰度分类和MUSE框架评估代理在不同清晰度下的表现。

Comments 21 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07749 2026-02-10 cs.AI 57%

Geo-Code: A Code Framework for Reverse Code Generation from Geometric Images Based on Two-Stage Multi-Agent Evolution

Geo-Code: 一种基于双阶段多智能体演化的几何图像反向代码生成框架

Zhenyu Wu, Yanxi Long, Jian Li, Hua Huang

机构 * School of Artificial Intelligence, Beijing Normal University, Beijing, China(人工智能学院,北京师范大学,北京,中国)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 Geo-Code提出一种基于双阶段多智能体演化的几何图像反向代码生成框架,通过像素级锚定和度量驱动的代码演化提升几何重建精度和视觉一致性。

Comments ICML2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00181 2026-02-10 cs.CR cs.AI cs.LG 57%

CHAI: Command Hijacking against embodied AI

CHAI:针对具身AI的命令劫持

Luis Burbano, Diego Ortiz, Qi Sun, Siwei Yang, Haoqin Tu, Cihang Xie, Yinzhi Cao, Alvaro A Cardenas

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 CHAI是一种针对具身AI的物理环境间接提示注入攻击,通过嵌入误导性自然语言指令,利用多模态语言解释能力,有效提升了攻击性能,凸显了对传统对抗鲁棒性之外的防御需求。

Comments This work has been accepted for publication at the IEEE Conference on Secure and Trustworthy Machine Learning (SaTML). The final version will be available on IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07092 2026-02-10 cs.MA cs.AI 57%

Lemon Agent Technical Report

柠檬代理技术报告

Haipeng Jiang, Kailong Ren, Zimo Yin, Zhetao Sun, Xin Gan, Guangyi Lv, Ming He, Peng Wang, Congli Yin, Hong Pan, Changwen Zhang, Shan Tong, Zhengyu Xu, Zeping Chen, Yubin Huangfu, Yanzhi Xu, Xing Su, Qin Feng, Dong An, Jianping Fan

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 柠檬代理是一种基于AgentCortex框架的多代理系统,通过自适应调度和三级上下文管理策略,优化资源利用和任务处理效率,实现复杂场景下的高准确率表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06653 2026-02-09 cs.RO cs.AI 57%

RAPID: Reconfigurable, Adaptive Platform for Iterative Design

RAPID: 可重构、自适应平台用于迭代设计

Zi Yin, Fanhong Li, Shurui Zheng, Jia Liu

机构 * Zi Yin Fanhong Li Shurui Zheng Jia Liu

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

AI总结 RAPID通过模块化硬件和软件栈实现机器人操作策略的快速迭代与多模态配置的高效实验

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03595 2026-02-09 cs.CV 57%

Refer-Agent: A Collaborative Multi-Agent System with Reasoning and Reflection for Referring Video Object Segmentation

Refer-Agent:一种具有推理和反思的协作多智能体系统用于指引用视频对象分割

Haichao Jiang, Tianming Liang, Wei-Shi Zheng, Jian-Fang Hu

机构 * Sun Yat-sen University(中山大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

AI总结 Refer-Agent通过协作多智能体系统结合推理与反思机制,提升指引用视频对象分割的性能与灵活性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02071 2026-02-09 cs.AI 57%

Human-AI Co-Embodied Intelligence for Scientific Experimentation and Manufacturing

人类-人工智能协同具身智能用于科学实验与制造

Xinyi Lin, Yuyang Zhang, Yuanhang Gan, Juntao Chen, Hao Shen, Yichun He, Lijun Li, Ze Yuan, Shuang Wang, Chaohao Wang, Rui Zhang, Na Li, Jia Liu

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 本研究提出人类-人工智能协同具身智能,通过智能体-物理实验系统提升科学实验与制造的精度与可扩展性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08641 2026-02-06 cs.AI q-fin.TR 57%

Resisting Manipulative Bots in Meme Coin Copy Trading: A Multi-Agent Approach with Chain-of-Thought Reasoning

抵制操纵机器人在表情包加密货币复制交易中的应用:一种基于多智能体的链式推理方法

Yichen Luo, Yebo Feng, Jiahua Xu, Yang Liu

机构 * UCL, Centre for Blockchain Technologies(伦敦大学区块链技术中心) The University of Hong Kong, FinTech Academy(香港大学金融科技学院) Nanyang Technological University(南洋理工大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

AI总结 本文提出一种基于多智能体和链式推理的复制交易系统,以抵御操纵机器人,通过多模态大语言模型提升预测准确度和经济表现,实现加密货币投资的稳健收益。

Journal ref Proceedings of the ACM Web Conference 2026 (WWW'26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21826 2026-02-06 cs.CL 57%

Mil-SCORE: Benchmarking Long-Context Geospatial Reasoning and Planning in Large Language Models

Mil-SCORE:大型语言模型中长上下文地理空间推理与规划的基准测试

Aadi Palnitkar, Mingyang Mao, Nicholas Waytowich, Vinicius G. Goecks, Xiaomin Lin

机构 * University of Maryland, College Park MD, USA(马里兰大学) ERA Lab, University of South Florida, Tampa FL, USA(佛罗里达大学埃拉实验室) DEVCOM Army Research Laboratory, Aberdeen Proving Ground MD, USA(国防部陆军研究实验室) EEHPC Lab, Johns Hopkins University, Baltimore MD, USA(约翰霍普金斯大学EEHPC实验室)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL

AI总结 Mil-SCORE是首个针对复杂军事规划情景的多跳问题数据集,旨在评估大型语言模型在长上下文地理空间推理与规划中的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10962 2026-02-06 cs.LG cs.AI 57%

WebSTAR: Scalable Data Synthesis for Computer Use Agents with Step-Level Filtering

WebSTAR: 可扩展的数据合成用于计算机使用代理的步级过滤

Yifei He, Pranit Chawla, Yaser Souri, Subhojit Som, Xia Song

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Microsoft(微软公司)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 WebSTAR通过步级过滤技术合成高质量数据,构建了WebSTAR和WebSCORE数据集,并训练了高效的过程奖励模型StepRM,提升计算机使用代理的训练效果和部署效率。

Comments Project website: https://yifei-he.github.io/webstar-website/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03376 2026-02-04 cs.RO cs.CV 57%

PlanTRansformer: Unified Prediction and Planning with Goal-conditioned Transformer

PlanTRansformer: 一种结合目标条件变换器的统一预测与规划

Constantin Selzer, Fabina B. Flohr

机构 * Department of Electrical Engineering and Information Technology, Intelligent Vehicles Lab (IVL), Munich University of Applied Science(电气工程与信息科技系,智能车辆实验室(IVL),慕尼黑应用科学大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 PlanTRansformer通过整合目标条件预测、动态可行性等技术,提升自动驾驶中的预测与规划性能。

Comments Submitted and accepted at IEEE IV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02961 2026-02-04 cs.AI 57%

Generative Engine Optimization: A VLM and Agent Framework for Pinterest Acquisition Growth

生成引擎优化:一个面向Pinterest获取增长的VLM和代理框架

Faye Zhang, Qianyu Cheng, Jasmine Wan, Vishwakarma Singh, Jinfeng Rao, Kofi Boakye

机构 * Stanford University(斯坦福大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 Pinterest提出GEO框架,通过微调VLM和AI代理预测用户搜索需求,构建语义连贯的集合页面,提升生成搜索时代的视觉平台流量增长。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00103 2026-02-03 cond-mat.soft cond-mat.mtrl-sci cs.AI 57%

Autonomous Multi-Agent AI for High-Throughput Polymer Informatics: From Property Prediction to Generative Design Across Synthetic and Bio-Polymers

自主多智能体AI用于高通量聚合物信息学:从性质预测到生成设计跨越合成与生物聚合物

Mahule Roy, Adib Bazgir, Arthur da Silva Sousa Santos, Yuwen Zhang

机构 * Institute of Biomedical Engineering University of Oxford(生物医学工程研究所牛津大学) Department of Mechanical and Aerospace Engineering University of Missouri–Columbia(机械与航空航天工程系密苏里大学-哥伦比亚分校) Center for Engineering, Modeling and Applied Social Sciences Federal University of ABC (UFABC)(工程、建模与应用社会科学中心巴西联邦大学ABC分校)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 本文提出一个自主多智能体AI系统,用于高通量聚合物信息学,实现从性质预测到生成设计的跨合成与生物聚合物应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22948 2026-02-02 cs.AI 57%

Alignment among Language, Vision and Action Representations

语言、视觉和动作表征的一致性

Nicola Milano, Stefano Nolfi

专题命中 多模态Agent :cross-modal(abstract);分类 cs.AI

AI总结 研究通过训练Transformer智能体执行自然语言指令,发现语言、视觉和动作表征在跨模态中呈现部分共享的语义结构,支持模态无关的语义组织。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04769 2026-02-02 cs.CV 57%

Vision-Language-Action (VLA) Models: Concepts, Progress, Applications and Challenges

视觉-语言-动作(VLA)模型:概念、进展、应用与挑战

Ranjan Sapkota, Yang Cao, Konstantinos I. Roumeliotis, Manoj Karkee

机构 * Cornell University(康奈尔大学) The Hong Kong University of Science and Technology(香港科学与技术大学) University of the Peloponnese(希腊皮洛斯大学)

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV

AI总结 本文综述了视觉-语言-动作(VLA)模型的概念、进展、应用与挑战,探讨了其在自动驾驶、医疗机器人等领域的应用及未来发展方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20831 2026-01-29 cs.AI cs.RO 57%

MemCtrl: Using MLLMs as Active Memory Controllers on Embodied Agents

MemCtrl: 使用 MLLMs 作为具身智能体的主动内存控制器

Vishnu Sashank Dorbala, Dinesh Manocha

机构 * University of Maryland, College Park(马里兰大学 College Park分校)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 MemCtrl 通过使用 MLLMs 的可训练内存头 μ 实现在线内存修剪,提升具身智能体在复杂指令下的任务完成能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20323 2026-01-29 cs.AI 57%

ECG-Agent: On-Device Tool-Calling Agent for ECG Multi-Turn Dialogue

ECG-Agent: 用于ECG多轮对话的设备端工具调用代理

Hyunseung Chung, Jungwoo Oh, Daeun Kyung, Jiho Kim, Yeonsu Kwon, Min-Gyu Kim, Edward Choi

机构 * KAIST(韩国科学技术院) Ajou University School of Medicine(全州大学医学院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 ECG-Agent是一种基于LLM的多轮ECG对话工具调用代理,通过ECG-MTD数据集实现了设备端高效且准确的ECG对话处理。

Comments Accepted to ICASSP 2026 (5 pages, 2 figures, 5 tables)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13251 2026-01-29 cs.AI cond-mat.mtrl-sci 57%

"DIVE" into Hydrogen Storage Materials Discovery with AI Agents

深入探索氢储存材料发现的AI智能体

Di Zhang, Xue Jia, Tran Ba Hung, Seong Hoon Jang, Linda Zhang, Ryuhei Sato, Yusuke Hashimoto, Toyoto Sato, Kiyoe Konno, Shin-ichi Orimo, Hao Li

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 DIVE通过多智能体工作流提升氢储存材料数据提取效率,实现快速逆向设计。

Comments 23 pages, 5 figures. The supplementary video is available at the GitHub link provided in the manuscript

Journal ref Chemical Science 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19507 2026-01-28 cs.CL 57%

Automated Safety Benchmarking: A Multi-agent Pipeline for LVLMs

自动化安全基准测试:一种用于大型视觉语言模型的多代理流程

Xiangyang Zhu, Yuan Tian, Zicheng Zhang, Qi Jia, Chunyi Li, Renrui Zhang, Heng Li, Zongrui Wang, Wei Sun

机构 * Shanghai AI Lab(上海人工智能实验室) PolyU HK(PolyU香港分校) East China Normal University(东华大学)

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CL

AI总结 VLSafetyBencher通过多代理流程自动化构建高质量安全基准,有效区分不同模型的安全性,提升LVLMs的安全评估效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18735 2026-01-27 cs.AI cs.LG 57%

Why Keep Your Doubts to Yourself? Trading Visual Uncertainties in Multi-Agent Bandit Systems

为何保留你的怀疑?多智能体老虎机系统中的视觉不确定性交易

Jusheng Zhang, Yijia Fan, Kaitong Cai, Jing Yang, Jiawei Yao, Jian Wang, Guanlong Qu, Ziliang Chen, Keze Wang

机构 * Sun Yat-sen University(中山大学) University of Washington(华盛顿大学) Snap Inc.(Snap公司) Syracuse University(雪城大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 Agora通过去中心化市场交易机制提升多智能体系统在视觉任务中的协调效率与经济性。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17920 2026-01-27 cs.AI 57%

Agentic AI for Self-Driving Laboratories in Soft Matter: Taxonomy, Benchmarks,and Open Challenges

代理AI在软物质中的自主实验室:分类、基准和开放挑战

Xuanzhou Chen, Audrey Wang, Stanley Yin, Hanyang Jiang, Dong Zhang

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

AI总结 本文探讨了自主实验室在软物质中的应用,提出了基于智能体的环境交互框架,回顾了闭环实验的主要方法,并指出了多模态表示、校准不确定性、安全探索和共享基准基础设施等开放挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01078 2026-01-26 cs.AI 57%

SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds

SimWorld:一种用于物理和社会世界中自主代理的开放式真实模拟器

Jiawei Ren, Yan Zhuang, Xiaokang Ye, Lingjun Mao, Xuhong He, Jianzhi Shen, Mrinaal Dogra, Yiming Liang, Ruixuan Zhang, Tianai Yue, Yiqing Yang, Eric Liu, Ryan Wu, Kevin Benavente, Rajiv Mandya Nagaraju, Muhammad Faayez, Xiyan Zhang, Dhruv Vivek Sharma, Xianrui Zhong, Ziqiao Ma, Tianmin Shu, Zhiting Hu, Lianhui Qin

机构 * UCSD(加州大学圣地亚哥分校) UVA(弗吉尼亚大学) UIUC(伊利诺伊大学香槟分校) JHU(约翰·霍普金斯大学) Purdue(Purdue 大学) PolyU USC(美国南加州大学) UMich(密歇根大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 SimWorld是一个基于Unreal Engine 5构建的开放式真实模拟器,旨在开发和评估LLM/VLM代理在复杂物理和社会环境中的能力,通过多代理配送任务验证其推理模式与局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14274 2026-01-23 cs.CV cs.LG math.AT 57%

TUN: Detecting Significant Points in Persistence Diagrams with Deep Learning

TUN:利用深度学习在持续图中检测显著点

Yu Chen, Hongwei Lin

机构 * School of Mathematical Sciences Zhejiang University Hangzhou, China(数学科学学院 浙江大学 杭州, 中国)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

AI总结 TUN通过深度学习方法在持续图中自动检测显著点,提升拓扑数据分析的实用性与可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏