arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

IEEE TPAMI

IEEE Transactions on Pattern Analysis and Machine Intelligence · 期刊 · Computer Vision

共收录 1562
2207.06400 2026-01-27 cs.CV

PyMAF-X: Towards Well-aligned Full-body Model Regression from Monocular Images

PyMAF-X:迈向从单目图像中恢复对齐的完整身体模型

Hongwen Zhang, Yating Tian, Yuxiang Zhang, Mengcheng Li, Liang An, Zhenan Sun, Yebin Liu

机构 * Department of Automation, Tsinghua University(自动化系,清华大学) Department of Computer Science and Technology, Nanjing University(计算机科学与技术系,南京大学) Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院)

AI总结 PyMAF-X通过改进的回归方法实现从单目图像中准确恢复对齐的完整身体模型,提升了网格与图像的对齐性能。

Comments Article in IEEE TPAMI 2023, Update project page: https://zhanghongwen.cn/pymaf-x, An eXpressive extension of PyMAF [arXiv:2103.16507] for monocular human/hand/face/whole-body motion capture

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12795 2026-01-21 cs.CV

Combating Noisy Labels through Fostering Self- and Neighbor-Consistency

通过促进自一致性与邻居一致性来对抗噪声标签

Zeren Sun, Yazhou Yao, Tongliang Liu, Zechao Li, Fumin Shen, Jinhui Tang

机构 * School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院) State Key Laboratory of Intelligent Manufacturing of Advanced Construction Machinery(先进施工机械智能制造国家重点实验室) School of Computer Science, Faculty of Engineering, the University of Sydney(悉尼大学工程学院计算机科学系) School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院)

AI总结 本文提出Jo-SNC方法,通过自一致性与邻居一致性提升模型对噪声标签的鲁棒性,结合自适应阈值和三元组正则化,有效识别并处理分布内和分布外噪声样本。

Comments accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16396 2026-01-21 cs.CV

VRP-UDF: Towards Unbiased Learning of Unsigned Distance Functions from Multi-view Images with Volume Rendering Priors

VRP-UDF: 向多视图图像中通过体积渲染先验实现无偏的无符号距离函数学习

Wenyuan Zhang, Chunsheng Wang, Kanle Shi, Yu-Shen Liu, Zhizhong Han

机构 * School of Software, Tsinghua University(清华大学软件学院) China Telecom Wanwei Information Technology Co., Ltd.(中国电信万维信息技术有限公司) Kuaishou Technology(快手技术) Department of Computer Science, Wayne State University(韦恩州立大学计算机科学系)

AI总结 VRP-UDF通过引入体积渲染先验,解决多视图图像中无偏学习无符号距离函数的问题,提升表面重建的鲁棒性和可扩展性。

Comments Accepted by TPAMI 2026 and ECCV 2024. Project page: https://wen-yuan-zhang.github.io/VolumeRenderingPriors/ . v1 is the conference version, and v2 is the journal extension version

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11350 2026-01-19 cs.LG cs.AI

FEATHer: Fourier-Efficient Adaptive Temporal Hierarchy Forecaster for Time-Series Forecasting

FEATHer:一种 Fourier 效率的自适应时间层次预测器用于时间序列预测

Jaehoon Lee, Seungwoo Lee, Younghwi Kim, Dohee Kim, Sunghyun Sim

AI总结 FEATHer 通过轻量级多尺度分解、共享密集时间内核和频率感知分支门控等方法,在受限边缘设备上实现了准确的长期时间序列预测。

Comments Submitted to IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.09126 2026-01-16 eess.IV cs.CV cs.LG

Learning Physics-Informed Noise Models from Dark Frames for Low-Light Raw Image Denoising

从暗帧中学习物理信息噪声模型以实现低光照RAW图像去噪

Hansen Feng, Lizhi Wang, Yiqi Huang, Yuzhi Wang, Lin Zhu, Hua Huang

机构 * School of Computer Science and Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学) School of Artificial Intelligence, Beijing Normal University(人工智能学院,北京师范大学) Engineering Research Center of Intelligent Technology and Educational Application, Ministry of Education(智能技术与教育应用工程研究中心,教育部) Megvii Technology(科大讯飞)

AI总结 本文提出从暗帧学习物理信息噪声模型,通过物理引导噪声解耦、物理感知代理模型和可微分布损失提升低光照RAW图像去噪性能。

Comments 18 pages, 13 figures. Accepted by IEEE TPAMI (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.12826 2026-01-16 cs.CV

Data-Driven Feature Tracking for Event Cameras With and Without Frames

基于数据驱动的事件相机特征跟踪(有和无帧)

Nico Messikommer, Carter Fang, Mathias Gehrig, Giovanni Cioffi, Davide Scaramuzza

机构 * Robotics and Perception Group(机器人与感知组) University of Zurich(苏黎世大学)

AI总结 本文提出了一种基于数据驱动的事件相机特征跟踪方法,通过结合事件和帧信息,提升在低延迟和低带宽下的特征跟踪性能。

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence (Volume: 47, Issue: 5, May 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09601 2026-01-15 cs.CV

Iterative Differential Entropy Minimization (IDEM) method for fine rigid pairwise 3D Point Cloud Registration: A Focus on the Metric

迭代微分熵最小化(IDEM)方法用于精细刚性点云配对3D点云配准:聚焦度量

Emmanuele Barberi, Felice Sfravara, Filippo Cucinotta

机构 * Department of Engineering, University of Messina(工程学院,墨西拿大学)

AI总结 IDEM方法通过微分熵度量提升点云配准的鲁棒性,有效应对密度差异、噪声和部分重叠等挑战。

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025, Available in IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06803 2026-01-15 cs.CV

DyDiT++: Diffusion Transformers with Timestep and Spatial Dynamics for Efficient Visual Generation

DyDiT++: 带时间步和空间动态的扩散变换器用于高效的视觉生成

Wangbo Zhao, Yizeng Han, Jiasheng Tang, Kai Wang, Hao Luo, Yibing Song, Gao Huang, Fan Wang, Yang You

机构 * National University of Singapore(新加坡国立大学) DAMO Academy, Alibaba Group(阿里云达摩院) Hupan Lab(虎扑实验室) Tsinghua University(清华大学)

AI总结 DyDiT++通过动态调整时间步和空间计算,提升视觉生成效率,减少计算成本并拓展应用范围。

Comments This paper was accepted to the IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) on January 9, 2026. arXiv admin note: substantial text overlap with arXiv:2410.03456

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09705 2026-01-13 cs.CV cs.AI cs.LG

Practical Continual Forgetting for Pre-trained Vision Models

实用的预训练视觉模型持续遗忘

Hongbo Zhao, Fei Zhu, Bolin Ni, Feng Zhu, Gaofeng Meng, Zhaoxiang Zhang

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science & Innovation, Chinese Academy of Sciences(人工智能与机器人中心,香港科学创新研究院,中国科学院) SenseTime Research(商汤科技研究院)

AI总结 该研究提出GS-LoRA方法,通过组稀疏正则化和原型信息监督,实现预训练视觉模型的持续遗忘,有效删除特定类别信息同时最小化对其他类别的影响。

Comments Accepted by TPAMI. arXiv admin note: substantial text overlap with arXiv:2403.11530

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07278 2026-01-13 cs.AI

Lifelong Learning of Large Language Model based Agents: A Roadmap

基于大语言模型代理的终身学习:路线图

Junhao Zheng, Chengming Shi, Xidi Cai, Qiuke Li, Duzhen Zhang, Chenxing Li, Dong Yu, Qianli Ma

AI总结 本文提出了一种基于大语言模型代理的终身学习框架,通过三个模块整合多模态输入、知识存储与动态环境交互,以实现持续适应和长期性能提升。

Comments Accepted to IEEE TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11514 2026-01-12 cs.CR cs.AI

Exploring the Vulnerabilities of Federated Learning: A Deep Dive into Gradient Inversion Attacks

探索联邦学习的漏洞:深入分析梯度反向攻击

Pengxin Guo, Runxi Wang, Shuang Zeng, Jinjing Zhu, Haoning Jiang, Yanran Wang, Yuyin Zhou, Feifei Wang, Hui Xiong, Liangqiong Qu

机构 * School of Computing and Data Science, The University of Hong Kong(计算与数据科学学院,香港大学) Department of Mathematics, The University of Hong Kong(数学系,香港大学) Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(人工智能推动学院,香港科学与技术大学(广州)) Department of Electronic and Electrical Engineering, Southern University of Science and Technology(电子与电气工程系,南方科技大学) Department of Biomedical Data Science, Stanford University(生物医学数据科学系,斯坦福大学) Department of Computer Science and Engineering, University of California, Santa Cruz(计算机科学与工程系,加州大学圣克鲁兹分校) Department of Electrical and Electronic Engineering, The University of Hong Kong(电子与电气工程系,香港大学) Materials Innovation Institute for Life Sciences and Energy (MILES), HKU-SIRI(生命科学与能源材料创新研究所(MILES),HKU-SIRI)

AI总结 本文系统分析了联邦学习中梯度反向攻击的三种类型,揭示了其性能、实用性及威胁因素,并提出三阶段防御策略以增强隐私保护。

Comments Accepted by IEEE TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11930 2026-01-12 cs.CV cs.AI

AtomThink: Multimodal Slow Thinking with Atomic Step Reasoning

AtomThink: 多模态慢思考与原子步骤推理

Kun Xiang, Zhili Liu, Terry Jingchen Zhang, Yinya Huang, Yunshuang Nie, Kaixin Cai, Yiyang Yin, Runhui Huang, Hanhui Li, Yihan Zeng, Yu-Jie Yuan, Jianhua Han, Lanqing Hong, Hang Xu, Xiaodan Liang

机构 * Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) ETH Zurich(苏黎世联邦理工学院) Hong Kong University of Science and Technology(香港科技大学) University of Hong Kong(香港大学) Noah’s Ark Lab(诺亚实验室) Yinwang Intelligent Technology Co., Ltd.(亿纬智能科技有限公司)

AI总结 AtomThink通过引入原子步骤推理,提升多模态大语言模型的推理性能和效率。

Comments TPAMI accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04279 2026-01-09 cs.CV

Controllable Generation with Text-to-Image Diffusion Models: A Survey

基于文本到图像扩散模型的可控生成:综述

Pu Cao, Feng Zhou, Qing Song, Lu Yang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

AI总结 本文综述了基于文本到图像扩散模型的可控生成方法,分析了不同条件下的生成机制及分类。

Comments TPAMI 2025; A collection of resources on controllable generation with text-to-image diffusion models: https://github.com/PRIV-Creation/Awesome-Controllable-T2I-Diffusion-Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05818 2026-01-06 cs.CV

LRANet++: Low-Rank Approximation Network for Accurate and Efficient Text Spotting

LRANet++:用于准确高效文本定位的低秩近似网络

Yuchen Su, Zhineng Chen, Yongkun Du, Zuxuan Wu, Hongtao Xie, Yu-Gang Jiang

机构 * College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院) Institute of Trustworthy Embodied AI, College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学可信具身人工智能研究院) School of Information Science and Technology, University of Science and Technology of China(中国科学技术大学信息科学与技术学院)

AI总结 LRANet++通过低秩近似和三重分配方案,实现高效准确的任意形状文本定位。

Comments Accepted by IEEE TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07330 2026-01-05 cs.CV cs.AI cs.SE

Revisiting Out-of-Distribution Detection in Real-time Object Detection: From Benchmark Pitfalls to a New Mitigation Paradigm

重新审视实时目标检测中的分布外检测:从基准陷阱到新的缓解范式

Changshun Wu, Weicheng He, Chih-Hong Cheng, Xiaowei Huang, Saddek Bensalem

AI总结 本文提出了一种新的训练时缓解范式,通过合成数据集减少目标检测中的分布外幻觉问题,有效降低YOLO模型的错误率。

Comments Accepted at IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09973 2026-01-01 cs.CV

Beyond Degradation Redundancy: Contrastive Prompt Learning for All-in-One Image Restoration

超越退化冗余:面向所有图像恢复的对比提示学习

Gang Wu, Junjun Jiang, Kui Jiang, Xianming Liu, Liqiang Nie

机构 * School of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术学院,哈尔滨工业大学) School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen)(计算机科学与技术学院,哈尔滨工业大学(深圳))

AI总结 本文提出对比提示学习框架,通过稀疏提示模块和对比提示正则化提升提示-任务对齐,实现更有效的图像恢复。

Comments Accepted by IEEE TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08243 2026-01-01 cs.CV

Hierarchical Context Alignment with Disentangled Geometric and Temporal Modeling for Semantic Occupancy Prediction

具有解耦几何和时间建模的层次上下文对齐用于语义占用预测

Bohan Li, Jiajun Deng, Yasheng Sun, Xiaofeng Wang, Xin Jin, Wenjun Zeng

机构 * Shanghai Jiao Tong University(上海交通大学) Eastern Institute of Technology(东部技术研究所) Ningbo Institute of Digital Twin(宁波数字孪生研究所) University of Adelaide(阿德莱德大学) Tokyo Institute of Technology(东京技术大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

AI总结 本文提出Hi-SOP方法,通过解耦几何和时间上下文进行层次化对齐,提升语义占用预测的准确性,在多个数据集上取得优异效果。

Comments IEEE TPAMI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18113 2026-01-01 cs.CV

Chrono: A Simple Blueprint for Representing Time in MLLMs

Chrono: 一种用于多模态大语言模型中表示时间的简单蓝图

Hector Rodriguez, Boris Meinardus, Anil Batra, Anna Rohrbach, Marcus Rohrbach

AI总结 Chrono提出了一种简单通用的序列蓝图,用于提升多模态大语言模型在视频时间定位和 grounded 视频问答任务中的性能。

Comments Code: https://github.com/sudo-Boris/mr-Blip. Submitted to IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI). Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22931 2025-12-30 cs.AI cs.LG

Geometric Structural Knowledge Graph Foundation Model

几何结构知识图谱基础模型

Ling Xin, Mojtaba Nayyeri, Zahra Makki Nayeri, Steffen Staab

机构 * University of Stuttgart(斯图加特大学) University of Southampton(南安普顿大学) Shahrood University of Technology(沙霍尔德大学)

AI总结 Gamma通过引入多头几何注意力机制,提升知识图谱推理的表达能力,优于现有方法Ultra,在零样本归纳链接预测中表现更优。

Comments Submitted to IEEE TPAMI, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12019 2025-12-22 cs.CV

LN3DIFF++: Scalable Latent Neural Fields Diffusion for Speedy 3D Generation

LN3DIFF++: 可扩展的潜在神经场扩散用于快速3D生成

Yushi Lan, Fangzhou Hong, Shangchen Zhou, Shuai Yang, Xuyi Meng, Yongwei Chen, Zhaoyang Lyu, Bo Dai, Xingang Pan, Chen Change Loy

AI总结 LN3DIFF++通过3D感知架构和变分自编码器实现快速高质量的3D生成,优于现有方法在推理速度和生成质量上。

Comments TPAMI 2025 version of LN3Diff. Project webpage: https://nirvanalan.github.io/projects/ln3diff/ Code: https://github.com/NIRVANALAN/LN3Diff

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14087 2025-12-17 cs.CV

GaussianPlant: Structure-aligned Gaussian Splatting for 3D Reconstruction of Plants

GaussianPlant: 基于结构对齐的高斯点云用于植物3D重建

Yang Yang, Risa Shinoda, Hiroaki Santo, Fumio Okura

机构 * The University of Osaka(大阪大学)

AI总结 GaussianPlant通过结构对齐的高斯点云方法,实现植物高精度外观与结构的3D重建,适用于植物表型分析。

Comments Submitted to IEEE TPAMI, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.16397 2025-12-16 math.OC

Augmenting Iterative Trajectory for Bilevel Optimization: Methodology, Analysis and Extensions

增强迭代轨迹以解决双层优化:方法、分析与扩展

Risheng Liu, Yaohua Liu, Shangzhi Zeng, Jin Zhang

AI总结 本文提出增强迭代轨迹(AIT)方法,通过改进初始化和轨迹截断技术,解决双层优化中收敛性问题,并在不同场景下验证其有效性。

Comments 18 pages. Accepted by IEEE TPAMI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06490 2025-12-15 eess.AS cs.AI cs.MM cs.SD eess.SP

Recent Advances in Discrete Speech Tokens: A Review

最近离散语音标记的进展:综述

Yiwei Guo, Zhihan Li, Hankun Wang, Bohan Li, Chongtian Shao, Hanglei Zhang, Chenpeng Du, Xie Chen, Shujie Liu, Kai Yu

机构 * MoE Key Lab of Artificial Intelligence, Jiangsu Key Lab of Language Computing(MoE人工智能关键实验室、江苏语言计算重点实验室;X-LANCE实验室、计算机科学与工程系、上海交通大学) X-LANCE Lab, Department of Computer Science and Engineering, Shanghai Jiao Tong University

AI总结 本文综述了离散语音标记的最新进展,分析了声音标记和语义标记的分类、方法及挑战,为未来研究提供方向。

Comments 26 pages, 8 figures, 3 tables. Accepted to IEEE TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10945 2025-12-13 cs.CV

MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation

MeViS:一种多模态数据集,用于指称运动表达视频分割

Henghui Ding, Chang Liu, Shuting He, Kaining Ying, Xudong Jiang, Chen Change Loy, Yu-Gang Jiang

机构 * Fudan University(复旦大学) Shanghai University of Finance and Economics(上海财经大学) Nanyang Technological University(南洋理工大学)

AI总结 MeViS数据集通过多模态数据提升视频分割中运动表达的理解能力,提出LMPM++方法实现新突破。

Comments IEEE TPAMI, Project Page: https://henghuiding.com/MeViS/

Journal ref in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 12, pp. 11400-11416, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10258 2025-12-12 cs.LG

R^2-HGP: A Double-Regularized Gaussian Process for Heterogeneous Transfer Learning

R²-HGP:一种双正则化高斯过程用于异构迁移学习

Duo Wang, Xinming Wang, Chao Wang, Xiaowei Yue, Jianguo Wu

机构 * Department of Control Science and Systems Engineering, Peking University(控制科学与系统工程系,北京大学) China Mobile Information Technology Co., Ltd.(中国移动信息科技有限公司) Department of Industrial and Systems Engineering, University of Iowa(工业与系统工程系,爱荷华大学) Department of Industrial Engineering, Institute for Quality and Reliability, Tsinghua University(工业工程系,质量与可靠性研究院,清华大学)

AI总结 本文提出R²-HGP框架,通过双正则化和物理知识整合,解决异构迁移学习中输入空间异构、先验知识忽略和负迁移问题。

Comments 17 pages, 9 figures. Under review for IEEE TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04519 2025-12-12 cs.CV

l0-Regularized Sparse Coding-based Interpretable Network for Multi-Modal Image Fusion

基于l0正则化稀疏编码的可解释网络用于多模态图像融合

Gargi Panda, Soumitra Kundu, Saumik Bhattacharya, Aurobinda Routray

机构 * Department of EE, IIT Kharagpur, India(印度IIT Kharagpur电子工程系) Rekhi Centre of Excellence for the Science of Happiness, IIT Kharagpur, India(印度IIT Kharagpur幸福科学卓越中心) Department of E&ECE, IIT Kharagpur, India(印度IIT Kharagpur电子工程与电子学系)

AI总结 本文提出基于l0正则化稀疏编码的可解释网络FNet,用于多模态图像融合,通过分离独特和共同特征提升融合质量并增强下游任务性能。

Comments Accetped by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09425 2025-12-11 eess.IV

QSMnet-INR: Single-Orientation Quantitative Susceptibility Mapping via Implicit Neural Representation in k-Space

通过k空间域隐式神经表示实现单方向定量磁 susceptibility 映射的QSMnet-INR

Xuan Cai, Ruo-Mi Guo, Xiao-Wen Luo, Jing Zhao, Silun Wang, Tao Tan, Yue Liu, Hongbin Han, Mengting Liu

AI总结 QSMnet-INR通过隐式神经表示和物理一致性损失,在k空间域中实现单方向QSM的高精度反演,提升结构恢复和伪影抑制能力。

Comments 14 pages, 12 figures; submitted to IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19912 2025-12-09 cs.CV cs.LG cs.RO

Enhanced Spatiotemporal Consistency for Image-to-LiDAR Data Pretraining

增强的时空一致性用于图像到LiDAR数据预训练

Xiang Xu, Lingdong Kong, Hui Shuai, Wenwei Zhang, Liang Pan, Kai Chen, Ziwei Liu, Qingshan Liu

机构 * College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院) School of Computing, Department of Computer Science, National University of Singapore(新加坡国立大学计算机学院) School of Computer Science, Nanjing University of Posts and Telecommunications(南京邮电大学计算机学院) Shanghai AI Laboratory(上海人工智能实验室) S-Lab, Nanyang Technological University(南洋理工大学S实验室)

AI总结 SuperFlow++通过整合时空线索提升图像到LiDAR数据预训练效果,实现更鲁棒的特征表示和更高效的自动驾驶感知。

Comments IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18405 2025-12-09 cs.CV cs.LG

Iwin Transformer: Hierarchical Vision Transformer using Interleaved Windows

Iwin Transformer:基于交错窗口的分层视觉Transformer

Simin Huo, Ning Li

AI总结 Iwin Transformer通过交错窗口注意力和深度可分离卷积实现无位置嵌入的分层视觉Transformer,提升图像分类、语义分割和视频动作识别等任务的性能。

Comments 17 pages, 12 figures, Submitted to IEEE Transactions on Pattern Analysis and Machine Intelligence. Add additional video experiments

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.05258 2025-12-08 cs.CV cs.LG cs.RO

Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving

多模态数据高效3D场景理解用于自动驾驶

Lingdong Kong, Xiang Xu, Jiawei Ren, Wenwei Zhang, Liang Pan, Kai Chen, Wei Tsang Ooi, Ziwei Liu

机构 * WorldBench Team Project Lead(WorldBench团队项目负责人)

AI总结 LaserMix++通过多模态方法提升自动驾驶中LiDAR数据高效3D场景理解,以更少标注实现更高精度。

Comments IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏