arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

IEEE TPAMI

IEEE Transactions on Pattern Analysis and Machine Intelligence · 期刊 · Computer Vision

共收录 1591 篇
2510.24795 2026-09-29 cs.CV cs.AI cs.LG cs.RO 版本更新

A Survey on Efficient Vision-Language-Action Models

高效视觉-语言-动作模型的综述

Zhaoshu Yu, Bo Wang, Pengpeng Zeng, Haonan Zhang, Ji Zhang, Zheng Wang, Lianli Gao, Jingkuan Song, Nicu Sebe, Heng Tao Shen

机构 * School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院) ; School of Computing and Artificial Intelligence, Southwest Jiaotong University(西南交通大学计算机科学与人工智能学院) ; School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院) ; Department of Information Engineering and Computer Science, University of Trento(特伦托大学信息工程与计算机科学系)

AI总结 本文综述了高效视觉-语言-动作模型,系统分类了模型设计、训练和数据收集三个核心领域,总结了最新方法并提出了未来研究方向。

Comments Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI). 20 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.24564 2026-09-22 cs.CV 新提交

HyperCLIP++: Fine-tuning CLIP forOpen-vocabulary Semantic Segmentation in Hyperbolic Space

HyperCLIP++:在双曲空间中微调CLIP以实现开放词汇语义分割

Zelin Peng, Zhengqin Xu, Changsong Wen, Yu Huang, Yaoming Wang, Xiaokang Yang, Wei Shen

机构 * Shanghai Jiao Tong University(上海交通大学) ; Meituan(美团)

AI总结 针对CLIP微调提升开放词汇分割的现象,提出HyperCLIP++,通过双曲空间半径缩放对齐层级,仅微调5%参数即在三个基准上达到最先进性能。

Comments Accept by TPAMI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18028 2026-09-22 cs.CV 版本更新

MambaX: Image Super-Resolution with State Predictive Control

MambaX:基于状态预测控制的图像超分辨率

Chenyu Li, Danfeng Hong, Bing Zhang, Zhaojie Pan, Naoto Yokoya, Jocelyn Chanussot

机构 * Southeast University(东南大学) ; Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空信息研究所) ; College of Resources and Environment, University of Chinese Academy of Sciences(中国科学院大学资源与环境学院) ; School of Mathematics, Southeast University(东南大学数学学院) ; Department of Complexity Science and Engineering, Graduate School of Frontier Sciences, the University of Tokyo(东京大学前沿科学研究院复杂科学与工程部门) ; Univ. Grenoble Alpes, Inria, CNRS, Grenoble INP, LJK(格勒诺布尔阿尔卑斯大学、Inria、CNRS、Grenoble INP、LJK)

AI总结 MambaX通过动态状态预测控制和多模态融合方法提升图像超分辨率性能。

Comments Published in IEEE TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22920 2026-09-22 cs.CL cs.AI 版本更新

Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey

多模态大语言模型的离散标记化:综合综述

Jindong Li, Yali Fu, Jiahong Liu, Linxiao Cao, Wei Ji, Menglin Yang, Irwin King, Ming-Hsuan Yang

AI总结 该综述首次系统分类并分析面向大语言模型的离散标记化方法,涵盖8种VQ变体,探讨其算法原理、训练动态及在单模态与多模态系统中的集成,指出码本崩溃等挑战并展望动态量化等方向。

Comments Published in IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07044 2026-09-22 cs.CL cs.AI cs.CV 版本更新

Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning

Lingshu:面向统一多模态医学理解与推理的通才基础模型

Weiwen Xu, Hou Pong Chan, Long Li, Mahani Aljunied, Ruifeng Yuan, Jianyu Wang, Chenghao Xiao, Guizhen Chen, Chaoqun Liu, Zhaodonghui Li, Yu Sun, Junao Shen, Chaojun Wang, Jie Tan, Deli Zhao, Tingyang Xu, Hao Zhang, Yu Rong

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团)

AI总结 针对医学多模态大模型知识覆盖有限、易幻觉和缺乏推理能力的问题,提出数据整理流程和医学专用模型Lingshu,经多阶段训练及强化学习增强,在三大医学任务上优于现有开源模型。

Comments Accepted by TPAMI. Our webpage is https://alibaba-damo-academy.github.io/lingshu. Models and training data are available at https://huggingface.co/lingshu-medical-mllm

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14280 2026-09-22 cs.CV cs.LG 版本更新

CLIP-Powered Domain Generalization and Domain Adaptation: A Comprehensive Survey

CLIP驱动的域泛化与域适应:综合综述

Jindong Li, Yongguang Li, Yali Fu, Jiahong Liu, Yixin Liu, Menglin Yang, Irwin King

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) ; The Jilin University(吉林大学) ; Griffith University(格里菲斯大学) ; The Chinese University of Hong Kong(香港中文大学)

AI总结 本综述系统梳理CLIP在域泛化与域适应中的应用,提出统一分类法,分类方法并分析趋势,为构建鲁棒模型提供指导。

Comments Published in IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.19716 2026-09-18 cs.CV 新提交

GAPrompt++: Multi-Granular Geometry-Aware Point Cloud Prompt for 3D Vision Model

GAPrompt++:面向3D视觉模型的多粒度几何感知点云提示

Zixiang Ai, Zhenyu Cui, Yufei Guo, Wenwen Qiang, Lei Chen, Jiwen Lu, Jiahuan Zhou

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王选计算机研究所) ; Intelligent Science and Technology Academy of CASIC(中国航天科工集团智能科技研究院) ; Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) ; Department of Automation, Tsinghua University(清华大学自动化系)

AI总结 针对3D点云任务适配中现有提示方法忽略几何结构的问题,提出多粒度几何感知提示方法GAPrompt++,通过点位移、关键点提示和提示传播机制,以少于2%可训练参数超越全量微调,并构建新基准。

Comments Accepted by TPAMI 2026. Code at https://github.com/PKU-OV3-LAB/GAPromptPlus.git

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21235 2026-09-18 stat.ML cs.AI cs.CV 版本更新

Domain Elastic Transform: Bayesian Function Registration for High-Dimensional Scientific Data

领域弹性变换:用于高维科学数据的贝叶斯函数注册

Osamu Hirose, Emanuele Rodola

机构 * Institute of Science and Engineering, Kanazawa University(金泽大学科学与工程研究所) ; Department of Computer Science, Sapienza University of Rome(罗马大学计算机科学系)

AI总结 本文提出领域弹性变换(DET),一种无网格概率框架,用于高维科学数据的几何与功能对齐,通过将数据视为不规则域上的函数,实现高维信号直接注册,无需分箱,且在大规模数据集上表现优异。

Comments 18 pages, 16 figures. Published in IEEE TPAMI. v3 is identical to v2; only the publication information was updated

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.11443 2026-09-18 cs.CV

LuvHarris: A Practical Corner Detector for Event-cameras

LuvHarris:一种适用于事件相机的实用角点检测器

Arren Glover, Aiko Dinale, Leandro De Souza Rosa, Simeon Bamford, Chiara Bartolozzi

机构 * Istituto Italiano di Tecnologia(意大利技术研究院)

AI总结 针对现有事件相机角点检测方法精度或实时性不足的问题,提出LuvHarris方法,通过新型阈值序数事件表面和优化的Harris算法实现,速度达前沿方法2.6倍以上,兼顾高精度与实时性。

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence ( Volume: 44, Issue: 12, 01 December 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24100 2026-09-17 cs.CV 版本更新

Think Before You Move: Latent Motion Reasoning for Text-to-Motion Generation

在行动前思考:用于文本到动作生成的潜在动作推理

Yijie Qian, Juncheng Wang, Yuxiang Feng, Chao Xu, Wang Lu, Yang Liu, Baigui Sun, Yiqiang Chen, Yong Liu, Shujun Wang

机构 * Zhejiang University(浙江大学) ; Hong Kong Polytechnic University(香港理工大学) ; IROOTECH TECHNOLOGY(IROOTECH技术) ; Wolf 1069 b Lab, Sany Group(Wolf 1069 b实验室,三一集团) ; Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

AI总结 本文提出Latent Motion Reasoning框架,通过分阶段推理与执行提升文本到动作生成的语义与物理合理性。

Comments Accepted to TPAMI, Project Page: https://chenhaoqcdyq.github.io/LMR/

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.03017 2026-09-16 cs.CV

RealLiFe: Real-Time Light Field Reconstruction via Hierarchical Sparse Gradient Descent

RealLiFe:基于分层稀疏梯度下降的实时光场重建

Yijie Deng, Lei Han, Tianpeng Lin, Lin Li, Jinzhi Zhang, Lu Fang

机构 * Dept. of Electrical Engineering, Tsinghua University, Beijing, China(电子工程系,清华大学,北京,中国) ; Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院) ; Huawei Technologies(华为技术)

AI总结 针对稀疏视图输入的实时光场重建需求,提出基于分层稀疏梯度下降的RealLiFe方法,借助MPI稀疏流形特性实现高质量实时重建,速度远超离线方法且性能优于现有在线方法。

Comments Accepted by IEEE TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.15239 2026-09-15 cs.LG cs.AI physics.soc-ph 新提交

ProtoGuide: Prototype-Driven Guidance for Class-Conditional Graph Generation

ProtoGuide:类条件图生成的原型驱动引导

Salvatore Romano, Marco Grassia, Pietro Liò, Giuseppe Mangioni

机构 * University of Catania(卡塔尼亚大学) ; University of Cambridge(剑桥大学) ; University Campus Bio-Medico of Rome(罗马生物医学自由大学)

AI总结 ProtoGuide提出一种事后、与骨干无关的框架,通过原型评分和梯度注入实现离散图扩散模型的类条件引导,在多个数据集上显著提升分类准确率。

Comments Preprint. Submitted to IEEE Transactions on Pattern Analysis and Machine Intelligence. 35 pages, 2 figures, 33 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.12965 2026-09-14 cs.CV cs.AI 新提交

Generative Retrieval for Unsupervised Text-Based Person Search

生成式检索用于无监督文本行人搜索

Mang Ye, Yucheng Ji, Yang Bai, Min Cao, Siyuan Chai, Bo Du, Min Zhang

机构 * Soochow University(苏州大学) ; Wuhan University(武汉大学) ; Zhipu AI(智谱AI) ; Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

AI总结 本文提出GTR+两阶段生成-检索框架,通过分层描述生成与自适应置信度加权学习实现无监督文本行人搜索,并贡献LargeFine-Person数据集,在多个基准上验证了有效性。

Comments 17 pages, 10 figures. Accepted for publication in IEEE Transactions on Pattern Analysis and Machine Intelligence

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 1-17, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.11875 2026-09-11 cs.RO 新提交

UniMPA: A Unified Memory-Prediction-Action Model via Action-Grounded Transition Modeling

UniMPA:一种通过动作接地转移建模的统一记忆-预测-动作模型

Wei Li, Rui Shao, Jie He, Lingsen Zhang, Ziwei Liu, Liqiang Nie

机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) ; Nanyang Technological University(南洋理工大学)

AI总结 针对VLA模型中的转移可实现性差距问题,提出UniMPA统一模型,通过动作接地转移接口、持久选择性预测和双记忆库检索,解决转移模糊性、预测执行不匹配及经验实现不匹配。

Comments Submitted to IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI). Project page: https://JiuTian-VL.github.io/UniMPA-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13160 2026-09-11 cs.LG cs.AI cs.CR cs.CV 版本更新

CertDW: Towards Certified Dataset Ownership Verification via Conformal Calibration

CertDW:通过保形校准实现认证的数据集所有权验证

Ting Qiao, Yiming Li, Jianbin Li, Yingjia Wang, Leyi Qi, Junfeng Guo, Ruili Feng, Dacheng Tao

机构 * School of Control and Computer Engineering, North China Electric Power University(控制与计算机工程学院,华北电力大学) ; College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学) ; School of Cybersecurity, Northwestern Polytechnical University(网络安全学院,西北工业大学) ; Department of Computer Science, University of Maryland(计算机科学系,马里兰大学) ; Alibaba Group(阿里巴巴集团)

AI总结 针对现有数据集所有权验证在扰动下失效的问题,提出首个认证水印CertDW,结合保形校准的统计度量,实现受攻击下可靠的所有权验证。

Comments To appear in TPAMI 2026. 28 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.07525 2026-09-09 cs.CV 新提交

When Semantically Consistent Encoding Meets View-Label Heterogeneity Modeling: A Unified Framework for Incomplete Multi-View Multi-Label Learning

当语义一致编码遇见视图-标签异质性建模:不完整多视图多标签学习的统一框架

Chengliang Liu, Bo Li, Bob Zhang, Yanghao Zhou, Jie Wen, Wenwu Wang

机构 * University of Macau(澳门大学) ; Beijing Institute of Technology(北京理工大学) ; Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) ; University of Surrey(萨里大学)

AI总结 本文提出V2L统一框架,通过扰动感知编码和主动视图-标签相关性建模,解决不完整多视图多标签学习中的语义一致与视图异质性难题,在五个基准上取得领先性能。

Comments Accepted by IEEE TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.06578 2026-09-09 cs.CV 新提交

Learning to Use Imagination: Progress-Conditioned Future Utilization for World Action Models

学习使用想象力:面向世界动作模型的进度条件化未来利用

Yijie Zhu, Zitong Yu, Wei Li, Hui Ma, Wen Li, Rui Shao, Liqiang Nie

机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) ; Great Bay University(大湾区大学) ; University of Electronic Science and Technology of China(电子科技大学)

AI总结 提出ProWAM,一种进度条件化世界动作模型,通过自监督双时间进度编码器和层次化进度条件化想象调制,自适应利用未来想象,提升VLA/WAM模型性能。

Comments Submitted to IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI). Project page: https://github.com/JiuTian-VL/ProWAM

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24131 2026-09-07 cs.LG cs.CV 版本更新

Reservoir-Based Graph Convolutional Networks

基于回声库的图卷积网络

Mayssa Soussia, Gita Ayu Salsabila, Mohamed Ali Mahjoub, Islem Rekik

机构 * National Engineering School of Sousse, University of Sousse, LATIS – Laboratory of Advanced Technology and Intelligent Systems(苏塞国家工程学校,苏塞大学,先进技术和智能系统实验室) ; BASIRA Lab, Imperial-X and Department of Computing, Imperial College London(BASIRA实验室,Imperial-X和伦敦帝国理工学院计算系)

AI总结 本文提出RGC-Net,通过整合回声库动态与结构化图卷积,解决GCN在复杂动态数据中的长距离依赖捕捉与过平滑问题,实现图分类和生成任务的高性能表现。

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06002 2026-09-07 cs.LG 版本更新

DeltaGNN: Graph Neural Network with Information Flow Control

DeltaGNN:具备信息流控制的图神经网络

Kevin Mancini, Islem Rekik

AI总结 针对图神经网络的过平滑、过压缩及长程交互检测难题,提出带信息流控制的DeltaGNN,在10类真实世界图数据集上验证了其可扩展、可泛化的优越性能。

Journal ref K. Mancini and I. Rekik, "DeltaGNN: Graph Neural Network With Information Flow Control," in IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.03117 2026-09-04 cs.LG cs.CV 新提交

Kernel Reboot: Breaking the Boundaries of Neural Tangent Kernels for Neural Fields

内核重启:突破神经场的神经正切核边界

Amir Mallak, Alaa Maalouf, Lior Wolf, Daniela Rus, Dan Rosenbaum

机构 * University of Haifa(海法大学) ; Massachusetts Institute of Technology(麻省理工学院) ; Tel Aviv University(特拉维夫大学)

AI总结 该研究针对神经场从稀疏观测重建的难题,提出NTK-KIP、MetaQuill、MetaQuill-KIP三种算法,实现了兼具非线性与元可学习性的神经场,提升了极稀疏观测下的重建与补全性能。

Comments Published in IEEE TPAMI, vol. 48, no. 9, pp. 10940-10957, Sep. 2026. Author version adds related-work references and biography updates; Figures 12 and 13 were regenerated from the same locked hyperparameter sweep. Tabulated results, reported best points, scientific claims, and conclusions are unchanged

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 48, no. 9, pp. 10940-10957, Sep. 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16046 2026-09-04 physics.soc-ph cs.CY cs.SI 版本更新

CARDIO-Affect: A Hamiltonian-Variability Framework for Spatio-Temporal Emotional Pattern Recognition with Manifold-Based Individual and Group Profiling

CARDIO-Affect:一种基于哈密顿变异性框架的时空情感模式识别方法,结合基于流形的个体和群体分析

Xiao Sun

AI总结 本文提出CARDIO-Affect框架,通过哈密顿变分理论分析长期情感动态,结合流形学习实现个体和群体情感识别,验证了复杂系统中情感的多稳定性、弱混沌等特征。

Comments v3: supersedes v2. Adds two-layer framework architecture figure (micro Langevin SDE <-> macro sparse network), 4 pillars, 45-D/18-D outputs. Companion: arXiv:2510.15221 (WELD). 23 pages. Submitted to IEEE TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.30839 2026-09-01 cs.CV 新提交

Physical Adversarial Examples for Person Detectors in Thermal Images Based on 3D Modeling

基于3D建模的热图像行人检测器的物理对抗样本

Xiaopei Zhu, Siyuan Huang, Zhanhao Hu, Jianmin Li, Jun Zhu, Xiaolin Hu

机构 * Tsinghua University(清华大学) ; Chinese Institute for Brain Research (CIBR)(中国脑科学研究院)

AI总结 该研究基于3D建模制作红外对抗服装,针对YOLOv9等热图像行人检测器实现高攻击成功率,且具有良好的可迁移性。

Comments Accepted by TPAMI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29230 2026-09-01 cs.CV 新提交

Compact Snapshot Spectral Imaging with Calibration-Free Aperture Diffraction

无校准孔径衍射的紧凑型快照光谱成像

Tao Lv, Quan Yuan, Shiqiao Li, Chenglong Huang, Linsen Chen, Chongde Zi, Shuming Wang, Xun Cao

机构 * Nanjing University(南京大学)

AI总结 该研究针对快照光谱成像系统复杂、需重复校准的问题,提出无校准的ADIS方法,结合ODAUVST框架实现紧凑型全分辨率光谱成像,验证了其性能优势。

Comments Submitted to IEEE TPAMI. Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13073 2026-09-01 cs.RO cs.CV 版本更新

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey

用于机器人操作的基于大型VLM的视觉-语言-动作模型:综述

Rui Shao, Wei Li, Lingsen Zhang, Renshan Zhang, Zhiyang Liu, Ran Chen, Liqiang Nie

机构 * School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen)(计算机科学与技术学院,哈尔滨工业大学(深圳))

AI总结 本综述首次系统分类梳理用于机器人操作的基于大型VLM的VLA模型,明确其定义与两类架构,考察其与先进领域的集成等内容,整合进展并提供更新项目页面

Comments Under Minor Revision at IEEE TPAMI, Project Page: https://github.com/JiuTian-VL/Large-VLM-based-VLA-for-Robotic-Manipulation

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26859 2026-08-28 cs.CV 新提交

A Geometry-Driven, Framework-Agnostic Optimization for Object Pose Estimation

面向物体姿态估计的几何驱动、框架无关优化方法

Wei Chen, Tao Zhen, Zhongchen Shi, Jing Zhang, Liang Xie, Erwei Yin

AI总结 该研究提出一种几何驱动、框架无关的物体姿态估计数据优化方法,通过主轴线对齐构建旋转表示,提升了姿态估计精度且无需修改网络架构。

Comments Submitted to TPAMI, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08897 2026-08-28 cs.CV cs.AI cs.CL cs.MM 版本更新

Recurrence Meets Transformers for Universal Multimodal Retrieval

循环机制与Transformer结合的通用多模态检索模型

Davide Caffagni, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara

机构 * Department of Education and Humanities, University of Modena and Reggio Emilia(教育与人文学院, Modena and Reggio Emilia大学)

AI总结 本文提出结合循环机制与Transformer的统一多模态检索模型ReT-2,其支持多模态查询,在M2KR等基准上达最优性能,推理更快、内存占用更低,还可提升下游任务表现。

Comments TPAMI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23330 2026-08-25 cs.CV 新提交

IntentQA: Intent Question Answering in Videos by Cognitive Context Reasoning

IntentQA:基于认知上下文推理的视频意图问答

Jiapeng Li, Ping Wei, Wenjuan Han, Song-Chun Zhu, Lifeng Fan

机构 * Beijing Institute for General Artificial Intelligence (BIGAI)(北京通用人工智能研究院) ; Beijing Jiaotong University(北京交通大学)

AI总结 本文提出视频意图问答任务 IntentQA,构建相关大规模数据集,提出 X-CaVIR 框架并引入对比性能下降指标,实验验证其有效性、优越性与稳定性。

Comments 18 pages, 7 figures. Accepted manuscript of an article published in IEEE Transactions on Pattern Analysis and Machine Intelligence

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 48, no. 9, pp. 11044-11061, September 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22679 2026-08-25 cs.CV cs.RO 新提交

Contextrast++: Robust Multi-Scale Contextual Contrastive Learning for Semantic Segmentation

Contextrast++:用于语义分割的鲁棒多尺度上下文对比学习方法

Changki Sung, Hyungtae Lim, Wanhee Kim, Youngwoo Seo, Hyun Myung

机构 * Information & Electronics Research Institute, KAIST(韩国科学技术院信息与电子研究院) ; Zoox Inc.(Zoox公司) ; KAIST (Korea Advanced Institute of Science and Technology)(韩国科学技术院) ; Hanwha Aerospace(韩华宇航) ; School of Electrical Engineering, KAIST(韩国科学技术院电气工程学院)

AI总结 针对语义分割中上下文捕捉与长尾分布问题,提出含CCL、BANE采样的Contextrast++,在无额外推理开销下提升了基于对比学习的SOTA方法性能。

Comments Accepted to IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26917 2026-08-25 cs.CV 版本更新

AnimateAnyMesh++: A Flexible Feed-Forward Framework for High-Fidelity Text-Driven Mesh Animation

AnimateAnyMesh++: 一种灵活的4D基础模型用于高质量文本驱动的网格动画

Zijie Wu, Chaohui Yu, Fan Wang, Xiang Bai

机构 * School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院) ; DAMO Academy, Alibaba Group(阿里巴巴达摩院) ; Hupan Lab, Hangzhou, China(湖畔实验室) ; School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院)

AI总结 本文提出AnimateAnyMesh++,通过扩展数据集、改进架构和生成能力,实现高质量文本驱动的网格动画,提升了轨迹重建和几何保真度。

Comments 15 pages, TPAMI 2026 accepted, code url: https://github.com/JarrentWu1031/AnimateAnyMesh-pp

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04001 2026-08-25 cs.CV 版本更新

Sa2VA: Marrying SAM2 with MLLM for Dense Grounded Understanding of Images and Videos

Sa2VA:将SAM2与多模态大语言模型结合用于图像与视频的密集接地理解

Haobo Yuan, Xiangtai Li, Tao Zhang, Yueyi Sun, Zilong Huang, Shilin Xu, Shunping Ji, Yunhai Tong, Lu Qi, Jiashi Feng, Ming-Hsuan Yang

机构 * University of California, Merced(加州大学默塞德分校) ; Bytedance Seed(字节跳动种子) ; Wuhan University(武汉大学) ; Peking University(北京大学)

AI总结 本研究提出Sa2VA,结合SAM2与MLLM实现图像视频密集接地理解,引入Ref-SAV数据集,在多任务中表现优异且可扩展至多款开源MLLM,代码模型已公开。

Comments Accepted by IEEE TPAMI. Code: https://github.com/Bytedance/Sa2VA

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
↑