arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

IEEE TPAMI

IEEE Transactions on Pattern Analysis and Machine Intelligence · 期刊 · Computer Vision

共收录 1575
2609.03117 2026-09-04 cs.LG cs.CV 新提交

Kernel Reboot: Breaking the Boundaries of Neural Tangent Kernels for Neural Fields

内核重启:突破神经场的神经正切核边界

Amir Mallak, Alaa Maalouf, Lior Wolf, Daniela Rus, Dan Rosenbaum

机构 * University of Haifa(海法大学) Massachusetts Institute of Technology(麻省理工学院) Tel Aviv University(特拉维夫大学)

AI总结 该研究针对神经场从稀疏观测重建的难题,提出NTK-KIP、MetaQuill、MetaQuill-KIP三种算法,实现了兼具非线性与元可学习性的神经场,提升了极稀疏观测下的重建与补全性能。

Comments Published in IEEE TPAMI, vol. 48, no. 9, pp. 10940-10957, Sep. 2026. Author version adds related-work references and biography updates; Figures 12 and 13 were regenerated from the same locked hyperparameter sweep. Tabulated results, reported best points, scientific claims, and conclusions are unchanged

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 48, no. 9, pp. 10940-10957, Sep. 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16046 2026-09-04 physics.soc-ph cs.CY cs.SI 版本更新

CARDIO-Affect: A Hamiltonian-Variability Framework for Spatio-Temporal Emotional Pattern Recognition with Manifold-Based Individual and Group Profiling

CARDIO-Affect:一种基于哈密顿变异性框架的时空情感模式识别方法,结合基于流形的个体和群体分析

Xiao Sun

AI总结 本文提出CARDIO-Affect框架,通过哈密顿变分理论分析长期情感动态,结合流形学习实现个体和群体情感识别,验证了复杂系统中情感的多稳定性、弱混沌等特征。

Comments v3: supersedes v2. Adds two-layer framework architecture figure (micro Langevin SDE <-> macro sparse network), 4 pillars, 45-D/18-D outputs. Companion: arXiv:2510.15221 (WELD). 23 pages. Submitted to IEEE TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.30839 2026-09-01 cs.CV 新提交

Physical Adversarial Examples for Person Detectors in Thermal Images Based on 3D Modeling

基于3D建模的热图像行人检测器的物理对抗样本

Xiaopei Zhu, Siyuan Huang, Zhanhao Hu, Jianmin Li, Jun Zhu, Xiaolin Hu

机构 * Tsinghua University(清华大学) Chinese Institute for Brain Research (CIBR)(中国脑科学研究院)

AI总结 该研究基于3D建模制作红外对抗服装,针对YOLOv9等热图像行人检测器实现高攻击成功率,且具有良好的可迁移性。

Comments Accepted by TPAMI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29230 2026-09-01 cs.CV 新提交

Compact Snapshot Spectral Imaging with Calibration-Free Aperture Diffraction

无校准孔径衍射的紧凑型快照光谱成像

Tao Lv, Quan Yuan, Shiqiao Li, Chenglong Huang, Linsen Chen, Chongde Zi, Shuming Wang, Xun Cao

机构 * Nanjing University(南京大学)

AI总结 该研究针对快照光谱成像系统复杂、需重复校准的问题,提出无校准的ADIS方法,结合ODAUVST框架实现紧凑型全分辨率光谱成像,验证了其性能优势。

Comments Submitted to IEEE TPAMI. Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13073 2026-09-01 cs.RO cs.CV 版本更新

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey

用于机器人操作的基于大型VLM的视觉-语言-动作模型:综述

Rui Shao, Wei Li, Lingsen Zhang, Renshan Zhang, Zhiyang Liu, Ran Chen, Liqiang Nie

机构 * School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen)(计算机科学与技术学院,哈尔滨工业大学(深圳))

AI总结 本综述首次系统分类梳理用于机器人操作的基于大型VLM的VLA模型,明确其定义与两类架构,考察其与先进领域的集成等内容,整合进展并提供更新项目页面

Comments Under Minor Revision at IEEE TPAMI, Project Page: https://github.com/JiuTian-VL/Large-VLM-based-VLA-for-Robotic-Manipulation

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26859 2026-08-28 cs.CV 新提交

A Geometry-Driven, Framework-Agnostic Optimization for Object Pose Estimation

面向物体姿态估计的几何驱动、框架无关优化方法

Wei Chen, Tao Zhen, Zhongchen Shi, Jing Zhang, Liang Xie, Erwei Yin

AI总结 该研究提出一种几何驱动、框架无关的物体姿态估计数据优化方法,通过主轴线对齐构建旋转表示,提升了姿态估计精度且无需修改网络架构。

Comments Submitted to TPAMI, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08897 2026-08-28 cs.CV cs.AI cs.CL cs.MM 版本更新

Recurrence Meets Transformers for Universal Multimodal Retrieval

循环机制与Transformer结合的通用多模态检索模型

Davide Caffagni, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara

机构 * Department of Education and Humanities, University of Modena and Reggio Emilia(教育与人文学院, Modena and Reggio Emilia大学)

AI总结 本文提出结合循环机制与Transformer的统一多模态检索模型ReT-2,其支持多模态查询,在M2KR等基准上达最优性能,推理更快、内存占用更低,还可提升下游任务表现。

Comments TPAMI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06002 2026-08-27 cs.LG 版本更新

DeltaGNN: Graph Neural Network with Information Flow Control

DeltaGNN:具备信息流控制的图神经网络

Kevin Mancini, Islem Rekik

AI总结 针对图神经网络的过平滑、过压缩及长程交互检测难题,提出带信息流控制的DeltaGNN,在10类真实世界图数据集上验证了其可扩展、可泛化的优越性能。

Journal ref K. Mancini and I. Rekik, "DeltaGNN: Graph Neural Network With Information Flow Control," in IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23330 2026-08-25 cs.CV 新提交

IntentQA: Intent Question Answering in Videos by Cognitive Context Reasoning

IntentQA:基于认知上下文推理的视频意图问答

Jiapeng Li, Ping Wei, Wenjuan Han, Song-Chun Zhu, Lifeng Fan

机构 * Beijing Institute for General Artificial Intelligence (BIGAI)(北京通用人工智能研究院) Beijing Jiaotong University(北京交通大学)

AI总结 本文提出视频意图问答任务 IntentQA,构建相关大规模数据集,提出 X-CaVIR 框架并引入对比性能下降指标,实验验证其有效性、优越性与稳定性。

Comments 18 pages, 7 figures. Accepted manuscript of an article published in IEEE Transactions on Pattern Analysis and Machine Intelligence

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 48, no. 9, pp. 11044-11061, September 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22679 2026-08-25 cs.CV cs.RO 新提交

Contextrast++: Robust Multi-Scale Contextual Contrastive Learning for Semantic Segmentation

Contextrast++:用于语义分割的鲁棒多尺度上下文对比学习方法

Changki Sung, Hyungtae Lim, Wanhee Kim, Youngwoo Seo, Hyun Myung

机构 * Information & Electronics Research Institute, KAIST(韩国科学技术院信息与电子研究院) Zoox Inc.(Zoox公司) KAIST (Korea Advanced Institute of Science and Technology)(韩国科学技术院) Hanwha Aerospace(韩华宇航) School of Electrical Engineering, KAIST(韩国科学技术院电气工程学院)

AI总结 针对语义分割中上下文捕捉与长尾分布问题,提出含CCL、BANE采样的Contextrast++,在无额外推理开销下提升了基于对比学习的SOTA方法性能。

Comments Accepted to IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26917 2026-08-25 cs.CV 版本更新

AnimateAnyMesh++: A Flexible Feed-Forward Framework for High-Fidelity Text-Driven Mesh Animation

AnimateAnyMesh++: 一种灵活的4D基础模型用于高质量文本驱动的网格动画

Zijie Wu, Chaohui Yu, Fan Wang, Xiang Bai

机构 * School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院) DAMO Academy, Alibaba Group(阿里巴巴达摩院) Hupan Lab, Hangzhou, China(湖畔实验室) School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院)

AI总结 本文提出AnimateAnyMesh++,通过扩展数据集、改进架构和生成能力,实现高质量文本驱动的网格动画,提升了轨迹重建和几何保真度。

Comments 15 pages, TPAMI 2026 accepted, code url: https://github.com/JarrentWu1031/AnimateAnyMesh-pp

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04001 2026-08-25 cs.CV 版本更新

Sa2VA: Marrying SAM2 with MLLM for Dense Grounded Understanding of Images and Videos

Sa2VA:将SAM2与多模态大语言模型结合用于图像与视频的密集接地理解

Haobo Yuan, Xiangtai Li, Tao Zhang, Yueyi Sun, Zilong Huang, Shilin Xu, Shunping Ji, Yunhai Tong, Lu Qi, Jiashi Feng, Ming-Hsuan Yang

机构 * University of California, Merced(加州大学默塞德分校) Bytedance Seed(字节跳动种子) Wuhan University(武汉大学) Peking University(北京大学)

AI总结 本研究提出Sa2VA,结合SAM2与MLLM实现图像视频密集接地理解,引入Ref-SAV数据集,在多任务中表现优异且可扩展至多款开源MLLM,代码模型已公开。

Comments Accepted by IEEE TPAMI. Code: https://github.com/Bytedance/Sa2VA

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18882 2026-08-24 cs.LG cs.AI eess.IV q-bio.NC 版本更新

SPD Matrix Learning for Neuroimaging Analysis: Perspectives, Methods, and Challenges

神经影像分析中的SPD矩阵学习:视角、方法与挑战

Ce Ju, Reinmar Kobler, Antoine Collas, Motoaki Kawanabe, Cuntai Guan, Bertrand Thirion

机构 * Inria(法国国家信息与自动化研究所) CEA(法国国家原子能与替代能源委员会) Université Paris-Saclay(巴黎萨克雷大学) Advanced Telecommunications Research Institute International (ATR)(国际先进电信研究机构) RIKEN Artificial Intelligence Project(理化学研究所人工智能项目) College of Computing and Data Science (CCDS) and Centre of AI in Medicine (C-AIM)(计算与数据科学学院和医学人工智能中心)

AI总结 本文提出SPD矩阵学习作为神经影像分析的新方法,通过统一视角整合多种模态数据,结合几何统计与AI技术,解决传统分析与新兴AI范式之间的桥梁问题。

Comments 18 pages, 2 figures, 2 tables; This work was accepted for publication in IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) in 2026. Copyright may be transferred without notice, after which this version may no longer be accessible

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.02155 2026-08-24 cs.LG cs.AI cs.CV stat.ML

Guaranteed Tensor Recovery Fused Low-rankness and Smoothness

Hailin Wang, Jiangjun Peng, Wenjin Qin, Jianjun Wang, Deyu Meng

机构 * School of Mathematics and Statistics, Southwest University(西南大学数学与统计学院) Macau Institute of Systems Engineering, Macau University of Science and Technology(澳门科技大学澳门系统工程研究所)

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(9):10990-11007, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18306 2026-08-20 cs.CV 新提交

High-Flux Count-Free Single-Photon 3D Cameras

高通量无计数单光子3D相机

Kaustubh Sadekar, Vivek K Goyal, David Maier, Atul Ingle

机构 * Portland State University(波特兰州立大学) Boston University(波士顿大学)

AI总结 针对SPAD单光子相机的堆积失真与数据瓶颈问题,提出结合自由运行捕获和分析合成软件流水线的计算成像方法,可在宽光照范围可靠捕获场景信息,助力其应用于高通量场景。

Comments Presented at IEEE ICCP 2026 (Best Paper Award Winner). To appear in IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19504 2026-08-18 cs.LG cs.AI 版本更新

MoE-Enhanced Explainable Deep Manifold Transformation for Complex Data Embedding and Visualization

面向复杂数据嵌入与可视化的混合专家增强型可解释深度流形变换

Zelin Zang, Yuhao Wang, Jinlin Wu, Hong Liu, Yue Shen, Zhen Lei, Stan Z. Li

机构 * Centre for Artificial Intelligence and Robotics (CAIR), HKISI-CAS and Westlake University(人工智能与机器人中心(CAIR),HKISI-CAS和西湖大学) Westlake University(西湖大学) School of Information and Electrical Engineering, Hangzhou City University(信息与电气工程学院,杭州市大学) Academy of Edge Intelligence Hangzhou City University(边缘智能学院,杭州市大学) CAIR, HKISI-CAS(人工智能与机器人中心(CAIR),HKISI-CAS) State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), CASIA(多模态人工智能系统国家重点实验室(MAIS),CASIA) School of Artificial Intelligence, University of Chinese Academy of Sciences (UCAS)(人工智能学院,中国科学院大学(UCAS)) Ant Group(蚂蚁集团)

AI总结 针对降维中精度与可解释性的权衡难题,提出MoE增强的DMT-ME方法,结合几何感知双曲映射器与MoE模型,实验验证其在降维精度与可解释性上均表现优异。

Comments 17 pages, 15 figures, accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11290 2026-08-18 quant-ph cs.AI cs.LG 版本更新

ShadowNet for Data-Centric Quantum System Learning

面向以数据为中心的量子系统学习的ShadowNet

Yuxuan Du, Yibo Yang, Tongliang Liu, Zhouchen Lin, Bernard Ghanem, Dacheng Tao

机构 * Nanyang Technological University(南洋理工大学) School of Physical and Mathematical Sciences(物理与数学科学学院) School of Artificial Intelligence(人工智能学院) Shanghai Jiao Tong University(上海交通大学) King Abdullah University of Science and Technology(阿卜杜拉国王科技大学) School of Computer Science(计算机科学学院) University of Sydney(悉尼大学) State Key Lab of General AI(通用人工智能国家重点实验室) School of Intelligence Science and Technology(智能科学与技术学院) Peking University(北京大学) Institute of Biomedical Engineering(生物医学工程研究所) University of Oxford(牛津大学)

AI总结 针对大量子系统学习的维度灾难问题,提出结合神经网络与经典阴影优势的以数据为中心范式,构建ShadowNet模型,在60量子比特的量子态层析和直接保真度估计任务中验证了其有效性。

Comments Accepted to IEEE Transactions on Pattern Analysis and Machine Intelligence. 20 pages. 12 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13104 2026-08-14 cs.CV 新提交

Online Learning of Correspondences between Images

图像间对应关系的在线学习

Michael Felsberg, Fredrik Larsson, Johan Wiklund, Niclas Wadströmer, Jörgen Ahlberg

AI总结 该研究提出一种基于奈曼卡方散度的在线迭代学习方法,用于解决通用成像几何下图像序列间的点对应问题,算法实时运行且性能优于现有方法。

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, ISSN 0162-8828, E-ISSN 1939-3539, Vol. 35, no 1, p. 118-129

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12773 2026-08-14 cs.CV cs.LG eess.IV 新提交

CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers

CW-BASS v2:基于基础模型教师的半监督分割中感知饱和度的伪标签选择

Ebenezer Tarubinga

机构 * Ebenworks Systems(埃本沃克斯系统公司)

AI总结 CW-BASS v2是一种感知饱和度的伪标签选择方法,它针对DINOv2教师模型的置信饱和问题,结合预留校准等技术,在6个基准数据集上恢复UniMatch V2操作点并提升性能。

Comments Submitted to IEEE TPAMI. 22 pages, 11 figures, 17 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09986 2026-08-12 cs.AI cs.LG 新提交

MIDAS: Mutual Information Disentanglement with Uncertainty-Aware Fusion for Incomplete Multimodal Sentiment Analysis

MIDAS:面向不完整多模态情感分析的基于不确定性感知融合的互信息解缠方法

Yuhua Wen, Yingying Zhou, Qifei Li, Yingming Gao, Zhengqi Wen, Jianhua Tao, Ya Li

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Zhongguancun Academy(中关村学院) Tsinghua University(清华大学)

AI总结 本研究针对多模态情感分析中模态不完整的问题,提出MIDAS框架,通过互信息解缠与不确定性感知融合实现鲁棒的多模态表示,在多个数据集上取得优于基线的性能。

Comments Accepted to IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09438 2026-08-11 cs.CV 新提交

Unveiling the Secret of AdaLN-Zero in Diffusion Transformer

揭示扩散Transformer中AdaLN-Zero的秘密

Jie Zhu, Mingyu Ding, Boqiang Duan, Leye Wang, Jingdong Wang

机构 * Peking University(北京大学) UC Berkeley(加州大学伯克利分校) Baidu(百度)

AI总结 本研究探究扩散Transformer(DiT)中AdaLN-Zero性能优于AdaLN的原因,发现零初始化是关键要素,提出AdaLN-Gaussian初始化策略与SE-adaLN-Zero机制,经多数据集实验验证其有效性与泛化性。

Comments Accept by IEEE TPAMI 2026, camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18816 2026-08-10 cs.CV 版本更新

Grad-ECLIP: Gradient-based Visual and Textual Explanations for CLIP

Grad-ECLIP: 基于梯度的CLIP视觉与文本解释

Chenyang Zhao, Kun Wang, Janet H. Hsiao, Antoni B. Chan

机构 * Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系) Division of Social Science and Department of Computer Science & Engineering, Hong Kong University of Science & Technology(香港科学与技术大学社会科学学院及计算机科学与工程系) SenseTime Group Ltd(时光集团有限公司)

AI总结 本文提出Grad-ECLIP方法,通过分解CLIP编码器架构并分析匹配相似度与中间空间特征的关系,生成有效热图以解释CLIP匹配结果。通过通道和空间权重提升视觉解释质量,并通过定性定量评估验证其有效性。

Journal ref Zhao C, Wang K, Hsiao J H, et al. Grad-eclip: Gradient-based visual and textual explanations for clip[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.07039 2026-08-10 quant-ph cs.LG 版本更新

Quantum Generative Diffusion Model: A Fully Quantum-Mechanical Model for Generating Quantum State Ensemble

量子生成扩散模型:一种用于生成量子态系综的全量子力学模型

Chuangtao Chen, Qinglin Zhao, MengChu Zhou, Zhimin He, Zhili Sun, Haozhen Situ

机构 * Macau University of Science and Technology(澳门科技大学) New Jersey Institute of Technology(新泽西理工学院) Foshan University(佛山大学) University of Surrey(萨里大学) South China Agricultural University(华南农业大学)

AI总结 本研究提出全量子力学模型QGDM,基于量子信道理论构建正向与反向过程,在多种量子态生成任务中性能优于现有模型,为量子生成建模提供了新框架。

Comments 31 pages, 15 tables. Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence. The supplementary material is included at the end of the manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04949 2026-08-07 cs.CV cs.AI 版本更新

DeepForgeSeal: Latent Space-Driven Semi-Fragile Watermarking for Deepfake Detection Using Adversarial Reinforcement Learning

DeepForgeSeal:基于对抗强化学习的隐空间驱动半脆弱深度伪造检测水印技术

Tharindu Fernando, Clinton Fookes, Sridha Sridharan

机构 * The Signal Processing, Artificial Intelligence and Vision Technologies (SAIVT), Queensland University of Technology, Australia(信号处理、人工智能与视觉技术研究所(SAIVT),昆士兰理工大学)

AI总结 该研究针对深度伪造检测中现有水印方法鲁棒性与脆弱性难以平衡的问题,提出基于对抗强化学习的隐空间驱动半脆弱水印框架,在CelebA和CelebA-HQ数据集上性能优于现有最优方法

Comments Accepted for Publication in IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05042 2026-08-06 cs.RO 新提交

BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation

BridgeVLA++:一种面向三维操作的数据高效、可泛化且内存增强的视觉-语言-动作框架

Peiyan Li, Yuze Zhu, Yixiang Chen, Qisen Ma, Yuan Xu, Jiabing Yang, He Guan, Yan Huang, Hongtao Wu, Xiao Ma, Tao Kong, Liang Wang, Tieniu Tan

机构 * New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别国家重点实验室(NLPR)) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) FiveAges ByteDance Seed(字节跳动种子实验室)

AI总结 本研究提出内存增强的三维VLA框架BridgeVLA++,通过新增时空记忆架构,在保留原模型数据效率与泛化能力的同时,提升了记忆相关操作性能,且在多任务与真实平台上验证了其有效性。

Comments This work has been submitted to the IEEE TPAMI for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04106 2026-08-06 cs.CV eess.IV 新提交

LoRetta: A Foundation Model and Extensive Dataset for Global-Scale Remote Sensing Dense Image Matching

LoRetta:面向全球尺度遥感密集图像匹配的基础模型与大规模数据集

Siwei Yu, Han Guo, Zhenwei Shi, Zhengxia Zou

机构 * Beihang University(北京航空航天大学)

AI总结 针对全球尺度遥感密集图像匹配的挑战,本文提出结合可匹配性感知仿射定位与引导式密集配准的基础模型LoRetta,并构建含原生可匹配性标签的LEVIR-GM基准,实验表明其性能优于现有模型且迁移性良好。

Comments 17 pages, 12 figures, 6 tables. Submitted to IEEE Transactions on Pattern Analysis and Machine Intelligence. Project page: https://siweiyu.com/work/loretta/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20206 2026-08-05 cs.CV 版本更新

Toward Visual Grounding: A Survey

视觉定位:一项综述

Linhui Xiao, Xiaoshan Yang, Xiangyuan Lan, Yaowei Wang, Changsheng Xu

AI总结 本综述梳理视觉定位的发展与背景,总结近年进展与新挑战,定义规范研究设置,介绍相关数据集与应用,提出未来方向,是该领域最全面的综述,适合不同阶段研究者。

Comments Accepted by TPAMI 2025. We keep tracing related works at https://github.com/linhuixiao/Awesome-Visual-Grounding, article publication page: https://ieeexplore.ieee.org/abstract/document/11235566

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 48, no. 3, pp. 2749-2771, March 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08394 2026-08-05 cs.LG 版本更新

Adversarial Purification by Consistency-aware Latent Space Optimization on Data Manifolds

基于数据流形上感知一致性隐空间优化的对抗净化

Shuhai Zhang, Jiahao Yang, Hui Luo, Jie Chen, Li Wang, Feng Liu, Bo Han, Mingkui Tan

机构 * South China University of Technology(华南理工大学) Pazhou Laboratory(琶洲实验室) The University of Melbourne(墨尔本大学) Institute of Optics and Electronics, CAS(中国科学院光电技术研究所) Peking University(北京大学) Peng Cheng Laboratory(鹏城实验室) University of Texas at Arlington(德克萨斯大学阿灵顿分校) Hong Kong Baptist University(香港浸会大学)

AI总结 针对对抗净化易过度校正的问题,提出CMAP方法,通过在一致性模型隐空间优化向量恢复干净数据,在CIFAR-10和ImageNet-100上提升了对抗鲁棒性与自然准确率。

Comments Accepted at TPAMI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.04694 2026-08-05 cs.CV cs.AI 交叉投稿

Multi-Camera Trajectory Forecasting with Trajectory Tensors

基于轨迹张量的多摄像头轨迹预测

Olly Styles, Tanaya Guha, Victor Sanchez

AI总结 该研究针对多摄像头轨迹预测(MCTF)问题,提出轨迹张量技术及对应编码器-解码器模型,在含15个摄像头视角600小时视频的自建数据库上验证,其模型性能优于现有相关方法。

Comments To appear in IEEE Transactions on Pattern Analysis and Machine Intelligence (tPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01647 2026-07-30 cs.CV eess.IV 版本更新

Neural Network Assisted Lifting Steps For Improved Fully Scalable Lossy Image Compression in JPEG 2000

用于改进JPEG 2000中完全可扩展有损图像压缩的神经网络辅助提升步骤

Xinyue Li, Aous Naman, David Taubman

AI总结 该研究在JPEG 2000的小波变换中加入神经网络辅助提升步骤,经端到端训练后可在保留其可扩展性的同时,实现最高17.4%的平均BD码率节省,提升有损图像压缩性能。

Comments This work has been submitted to the IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏