arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

IEEE TPAMI

IEEE Transactions on Pattern Analysis and Machine Intelligence · 期刊 · Computer Vision

至 收录 20
2502.18816 2026-08-10 cs.CV 版本更新

Grad-ECLIP: Gradient-based Visual and Textual Explanations for CLIP

Grad-ECLIP: 基于梯度的CLIP视觉与文本解释

Chenyang Zhao, Kun Wang, Janet H. Hsiao, Antoni B. Chan

机构 * Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系) Division of Social Science and Department of Computer Science & Engineering, Hong Kong University of Science & Technology(香港科学与技术大学社会科学学院及计算机科学与工程系) SenseTime Group Ltd(时光集团有限公司)

AI总结 本文提出Grad-ECLIP方法,通过分解CLIP编码器架构并分析匹配相似度与中间空间特征的关系,生成有效热图以解释CLIP匹配结果。通过通道和空间权重提升视觉解释质量,并通过定性定量评估验证其有效性。

Journal ref Zhao C, Wang K, Hsiao J H, et al. Grad-eclip: Gradient-based visual and textual explanations for clip[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026

URL PDF HTML 收藏
2401.07039 2026-08-10 quant-ph cs.LG 版本更新

Quantum Generative Diffusion Model: A Fully Quantum-Mechanical Model for Generating Quantum State Ensemble

量子生成扩散模型:一种用于生成量子态系综的全量子力学模型

Chuangtao Chen, Qinglin Zhao, MengChu Zhou, Zhimin He, Zhili Sun, Haozhen Situ

AI总结 本研究提出全量子力学模型QGDM,基于量子信道理论构建正向与反向过程,在多种量子态生成任务中性能优于现有模型,为量子生成建模提供了新框架。

Comments 31 pages, 15 tables. Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence. The supplementary material is included at the end of the manuscript

URL PDF HTML 收藏
2511.04949 2026-08-07 cs.CV cs.AI 版本更新

DeepForgeSeal: Latent Space-Driven Semi-Fragile Watermarking for Deepfake Detection Using Adversarial Reinforcement Learning

DeepForgeSeal:基于对抗强化学习的隐空间驱动半脆弱深度伪造检测水印技术

Tharindu Fernando, Clinton Fookes, Sridha Sridharan

机构 * The Signal Processing, Artificial Intelligence and Vision Technologies (SAIVT), Queensland University of Technology, Australia(信号处理、人工智能与视觉技术研究所(SAIVT),昆士兰理工大学)

AI总结 该研究针对深度伪造检测中现有水印方法鲁棒性与脆弱性难以平衡的问题,提出基于对抗强化学习的隐空间驱动半脆弱水印框架,在CelebA和CelebA-HQ数据集上性能优于现有最优方法

Comments Accepted for Publication in IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

URL PDF HTML 收藏
2412.20206 2026-08-05 cs.CV 版本更新

Toward Visual Grounding: A Survey

视觉定位:一项综述

Linhui Xiao, Xiaoshan Yang, Xiangyuan Lan, Yaowei Wang, Changsheng Xu

AI总结 本综述梳理视觉定位的发展与背景,总结近年进展与新挑战,定义规范研究设置,介绍相关数据集与应用,提出未来方向,是该领域最全面的综述,适合不同阶段研究者。

Comments Accepted by TPAMI 2025. We keep tracing related works at https://github.com/linhuixiao/Awesome-Visual-Grounding, article publication page: https://ieeexplore.ieee.org/abstract/document/11235566

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 48, no. 3, pp. 2749-2771, March 2026

URL PDF HTML 收藏
2412.08394 2026-08-05 cs.LG 版本更新

Adversarial Purification by Consistency-aware Latent Space Optimization on Data Manifolds

基于数据流形上感知一致性隐空间优化的对抗净化

Shuhai Zhang, Jiahao Yang, Hui Luo, Jie Chen, Li Wang, Feng Liu, Bo Han, Mingkui Tan

AI总结 针对对抗净化易过度校正的问题,提出CMAP方法,通过在一致性模型隐空间优化向量恢复干净数据,在CIFAR-10和ImageNet-100上提升了对抗鲁棒性与自然准确率。

Comments Accepted at TPAMI 2026

URL PDF HTML 收藏
2606.14636 2026-07-30 cs.LG 版本更新

Projective Graph Residualization: Variation-Allocation Frontiers for Control-Function IV

用于控制函数工具变量的图扩散残差

Rui Wu, Zongyuan Chen, Hong Xie, Defu Lian, Enhong Chen

机构 * School of Computer Science and Engineering, University of Science and Technology of China(中国科学技术大学计算机科学与技术学院)

AI总结 提出自适应各向异性工具热流(A-IHF),一种基于图扩散的残差提取方法,用于灵活控制函数,通过检测处理跳跃并调整图传导性,在合成基准测试中优于多种基线方法。

Comments Submitted to IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI). 33 pages, 9 figures, including supplementary material

URL PDF HTML 收藏
2403.01647 2026-07-30 cs.CV eess.IV 版本更新

Neural Network Assisted Lifting Steps For Improved Fully Scalable Lossy Image Compression in JPEG 2000

用于改进JPEG 2000中完全可扩展有损图像压缩的神经网络辅助提升步骤

Xinyue Li, Aous Naman, David Taubman

AI总结 该研究在JPEG 2000的小波变换中加入神经网络辅助提升步骤,经端到端训练后可在保留其可扩展性的同时,实现最高17.4%的平均BD码率节省,提升有损图像压缩性能。

Comments This work has been submitted to the IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) for possible publication

URL PDF HTML 收藏
2411.19715 2026-07-27 cs.CV cs.CR cs.LG 版本更新

Forensics Adapter: Unleashing CLIP for Generalizable Face Forgery Detection

取证适配器:释放CLIP用于通用人脸伪造检测

Xinjie Cui, Yuezun Li, Delong Zhu, Jiaran Zhou, Junyu Dong, Siwei Lyu

AI总结 研究旨在将CLIP转变为人脸伪造检测器,提出取证适配器及扩展方法取证适配器++。通过适配器学习伪造痕迹,用交互策略增强视觉令牌,实现通用性。方法参数少性能优,扩展方法进一步提升性能,为相关检测提供基线。

Comments TPAMI 2026

URL PDF HTML 收藏
2506.00625 2026-07-21 cs.CV 版本更新

PI-H2T: Enhancing Long-Tailed Visual Recognition with Permutation-Invariant and Head-to-Tail Feature Fusion

PI-H2T:通过排列不变和头到尾特征融合增强长尾视觉识别

Mengke Li, Zhikai Hu, Yang Lu, Weichao Lan, Yiu-ming Cheung, Hui Huang

机构 * Guangdong Provincial Key Laboratory of Visual Media and Multidimensional Intelligence, CSSE, Shenzhen University, China(广东省视觉媒体与多维智能重点实验室,计算机科学与电子技术学院,深圳大学,中国) Fujian Key Laboratory of Sensing and Computing for Smart City, School of Informatics, Xiamen University, China(福建省智慧城市感知与计算重点实验室,信息学院,厦门大学,中国) Department of Computer Science, Hong Kong Baptist University, China(计算机科学系,香港 Baptist 大学,中国) Inspur Smart City Technology Co., Ltd., China(Inspur 智能城市技术有限公司,中国)

AI总结 针对长尾数据不平衡致深度学习模型准确率低的问题,提出PI-H2T方法,通过排列不变表示融合增强表示空间,利用头到尾融合调整分类器,经理论分析和实验验证其有效性,且可无缝集成到现有方法中提升性能。

Comments accpeted by TPAMI in 2026

Journal ref 10.1109/TPAMI.2026.3711911

URL PDF HTML 收藏
2408.10581 2026-07-16 cs.CV 版本更新

Multi-view Hand Reconstruction with a Point-Embedded Transformer

基于点嵌入变换器的多视图手部重建

Lixin Yang, Licheng Zhong, Pengxiang Zhu, Xinyu Zhan, Junxiao Kong, Jian Xu, Cewu Lu

机构 * School of Artificial Intelligence (SAI), Shanghai Jiao Tong University(人工智能学院(SAI),上海交通大学) School of Mechanical Engineering, Shanghai Jiao Tong University(机械工程学院,上海交通大学) School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University(电子信息与电气工程学院,上海交通大学) Institute of Automation Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所(CASIA))

AI总结 研究提出POEM模型用于多视图手部重建,通过在多视图立体空间嵌入基点表示手部网格,并结合多数据集及相机参数随机化训练,实现了通用、实用且经济高效的双手运动捕捉。

Comments TPAMI 2025, Extension of CVPR 2023, correction on Table 4: HO3D results

URL PDF HTML 收藏
2512.17788 2026-07-15 cs.LG 版本更新

Calibratable Disambiguation Loss for Multi-Instance Partial-Label Learning

用于多实例部分标签学习的可校准消歧损失

Wei Tang, Yin-Fang Yang, Weijia Zhang, Min-Ling Zhang

机构 * School of Computer Science and Engineering, Southeast University(计算机科学与工程学院,东南大学) Key Laboratory of Computer Network and Information Integration (Southeast University), MoE, China(计算机网络与信息集成重点实验室(东南大学),教育部,中国) School of Information and Physical Sciences, The University of Newcastle(信息与物理科学学院,新castle大学)

AI总结 针对多实例部分标签学习中校准不佳问题,提出可校准消歧损失(CDL),通过顶级与竞争对手预测边际调制消歧目标,有两个变体,经理论分析和实验验证,显著提升分类准确率和预期校准误差。

Comments Accepted at IEEE TPAMI. The code can be found at \url{https://github.com/tangw-seu/MIPLCDL}

URL PDF HTML 收藏
2607.08250 2026-07-14 cs.CV 版本更新

On the Design of Mixture-of-Experts for Dynamic Gaussian Splatting

关于动态高斯点云的专家混合模型设计

In-Hwan Jin, Hyeongju Mun, Joonsoo Kim, Kugjin Yun, Kyeongbo Kong

机构 * Pusan National University(釜山国立大学) Electronics and Telecommunications Research Institute (ETRI)(电子和电信研究所)

AI总结 研究动态3D高斯表示的多变形建模问题,提出变形专家混合(MoDE)和动态高斯点云专家混合(MoE-GS)两种方法,通过不同集成约束实现多变形建模,为该领域提供替代策略并阐明集成约束对变形专家设计及行为的影响。

Comments IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI, 2026)

URL PDF HTML 收藏
2403.01560 2026-07-10 cs.CV 版本更新

XOV-Action: Towards Generalizable Open-Vocabulary Action Recognition

XOV-Action:迈向可泛化的开放词汇动作识别

Kun-Yu Lin, Henghui Ding, Jia-Run Du, Jiaming Zhou, Yi-Xing Peng, Yu-Ming Tang, Zhilin Zhao, Chen Change Loy, Wei-Shi Zheng

机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) Institute of Big Data, Fudan University(复旦大学大数据研究院) AI Thrust, Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)人工智能方向) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院)

AI总结 研究针对开放词汇动作识别模型泛化性不足的问题,提出XOV-Action模型,通过学习多样化表示和场景无关表示克服两大挑战,还构建XOVABench基准,实验证明该模型能有效提升跨视频领域的动作识别性能。

Comments Accepted by TPAMI

URL PDF HTML 收藏
2408.00001 2026-07-08 cs.CV cs.AI cs.CY 版本更新

Replication in Visual Diffusion Models: A Survey and Outlook

视觉扩散模型中的复制:一项综述与展望

Wenhao Wang, Yifan Sun, Zongxin Yang, Zhengdong Hu, Zhentao Tan, Yi Yang

机构 * Australian Artificial Intelligence Institute, University of Technology Sydney(澳大利亚人工智能研究所,悉尼技术大学) Baidu Inc.(百度公司) College of Computer Science and Technology, Zhejiang University(计算机科学与技术学院,浙江大学)

AI总结 本文综述视觉扩散模型中的复制现象,将现有研究分类为揭示、理解和缓解该现象的方法,还回顾其现实影响,讨论了检测和基准测试复制的挑战及未来方向,助力研究人员和从业者理解AI技术与社会公益交叉点。

Comments Accepted by TPAMI 2026

URL PDF HTML 收藏
2106.15277 2026-07-07 cs.CV 版本更新

EPMF: Efficient Perception-aware Multi-sensor Fusion for 3D Semantic Segmentation

EPMF: 高效感知感知多传感器融合用于3D语义分割

Mingkui Tan, Zhuangwei Zhuang, Sitao Chen, Rong Li, Kui Jia, Qicheng Wang, Yuanqing Li

机构 * School of Software Engineering, South China University of Technology(南方科技大学软件工程学院) Pazhou Laboratory, Guangzhou, China(广州帕佐实验室) School of Electronic and Information Engineering, South China University of Technology(南方科技大学电子与信息工程学院) Minieye, Shenzhen, Guangdong, China(深圳Minieye公司)

AI总结 提出一种高效感知多传感器融合方法EPMF,通过透视投影对齐点云与RGB图像,利用残差融合模块和感知损失提升3D语义分割性能,在nuScenes上mIoU超越RangeFormer 0.9%。

Comments 16 pages, 12 figures, 14 tables, IEEE TPAMI 2024, extended version of the ICCV2021 paper

URL PDF HTML 收藏
2501.17015 2026-06-19 cs.AI cs.MA cs.RO 版本更新

UniMM: A Unified Mixture Model Framework for Multi-Agent Simulation

UniMM:一种用于多智能体仿真的统一混合模型框架

Longzhong Lin, Xuewu Lin, Kechun Xu, Haojian Lu, Lichao Huang, Rong Xiong, Yue Wang

机构 * Zhejiang University(浙江大学) Horizon Robotics

AI总结 提出UniMM框架统一回归混合模型与离散NTP模型,通过闭环样本生成缓解分布偏移,并在WOSAC基准上取得最优性能。

Comments Accepted author manuscript. The version of record has been published in IEEE Transactions on Pattern Analysis and Machine Intelligence

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, Early Access, 2026

URL PDF HTML 收藏
2410.15595 2026-06-18 cs.AI cs.CL cs.LG 版本更新

A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

直接偏好优化综述:数据集、理论、变体及应用

Wenyi Xiao, Zechuan Wang, Leilei Gan, Shuai Zhao, Zongrui Li, Ruirui Lei, Wanggui He, Luu Anh Tuan, Long Chen, Hao Jiang, Zhou Zhao, Fei Wu

机构 * Zhejiang University(浙江大学) Nanyang Technological University(南洋理工大学) Alibaba Group(阿里巴巴集团)

AI总结 综述直接偏好优化(DPO)在理论、变体、数据集和应用方面的进展,指出其作为RL-free替代方案的潜力与局限,并提出未来研究方向。

Comments Accepted by TPAMI 2026. Project page: https://github.com/Mr-Loevan/DPO-Survey

URL PDF HTML 收藏
2503.08038 2026-06-18 cs.LG cs.AI cs.CV 版本更新

Generalized Kullback-Leibler Divergence Loss

广义Kullback-Leibler散度损失

Jiequan Cui, Beier Zhu, Qingshan Xu, Zhuotao Tian, Xiaojuan Qi, Bei Yu, Hanwang Zhang, Richang Hong

机构 * Hefei University of Technology(合肥工业大学) University of Science and Technology of China(中国科学技术大学) Nanyang Technological University(南洋理工大学) The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

AI总结 本文提出广义KL散度损失,通过解耦KL损失为加权MSE和交叉熵损失,并引入非对称优化修正和类别全局信息,在对抗训练和知识蒸馏中取得SOTA性能。

Comments TPAMI 2026, extension of our NeurIPS paper "Decoupled Kullback-Leibler Divergence Loss". arXiv admin note: substantial text overlap with arXiv:2305.13948

URL PDF HTML 收藏
2508.09977 2026-06-16 cs.CV 版本更新

A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation

3D高斯泼溅应用综述:分割、编辑与生成

Shuting He, Peilin Ji, Yitong Yang, Changshuo Wang, Jiayi Ji, Yinglin Wang, Henghui Ding

机构 * Shanghai University of Finance and Economics(上海财经大学) University College London(伦敦大学学院) Xiamen University(厦门大学) Fudan University(复旦大学)

AI总结 综述3D高斯泼溅在分割、编辑和生成三大任务中的应用,总结代表性方法、监督策略和学习范式,并分析公共基准上的比较结果。

Comments IEEE TPAMI, GitHub Repo: https://github.com/heshuting555/Awesome-3DGS-Applications

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026

URL PDF HTML 收藏
2507.02288 2026-06-16 cs.CV cs.LG 版本更新

Prompt Disentanglement via Language Guidance and Representation Alignment for Domain Generalization

基于语言引导与表示对齐的提示解缠用于域泛化

De Cheng, Zhipeng Xu, Xinyang Jiang, Dongsheng Li, Nannan Wang, Xinbo Gao

机构 * School of Telecommunications Engineering, the State Key Laboratory of Integrated Services Networks (ISN), Xidian University, Xi’an, China(电信工程学院、集成服务网络国家重点实验室(ISN)、西安电子科技大学) Microsoft Research Asia, Shanghai, China(微软亚洲研究院,上海,中国)

AI总结 提出利用大语言模型自动解缠文本提示,并引入最差显式表示对齐,结合抽象提示增强源域多样性,实现域不变视觉表示学习,在多个基准上超越现有方法。

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 48, no. 6, pp. 6799-6816, June 2026

URL PDF HTML 收藏