arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

IEEE TPAMI

IEEE Transactions on Pattern Analysis and Machine Intelligence · 期刊 · Computer Vision

至 收录 1562
2403.04780 2026-05-26 cs.CL cs.AI

Graph-oriented Instruction Tuning of Large Language Models for Generic Graph Mining

面向通用图挖掘的大语言模型图导向指令微调

Yanchao Tan, Hang Lv, Pengxiang Zhan, Shiping Wang, Carl Yang

机构 * Engineering Research Center of Big Data Intelligence, Ministry of Education(教育部大数据智能工程研究中心) Fujian Key Laboratory of Network Computing and Intelligent Information Processing(福建省网络计算与智能信息处理重点实验室) College of Computer and Data Science, Fuzhou University(福州大学计算机与数据科学学院) Department of Computer Science, Emory University(埃默里大学计算机科学系)

AI总结 提出MuseGraph框架,通过紧凑图描述、基于思维链的指令生成和图感知指令微调,将GNN与LLM结合,实现跨任务和数据集的高效图挖掘。

Comments Accepted by TPAMI 2025

Journal ref IEEE Trans. Pattern Anal. Mach. Intell., vol. 48, no. 1, pp. 155-169, Jan. 2026

URL PDF HTML 收藏
2605.22423 2026-05-25 cs.CV

Moment-Reenacting: Inverse Motion Degradation with Cross-shutter Guidance

时刻重现:基于交叉快门引导的逆运动退化

Xiang Ji, Guixu Lin, Zhengwei Yin, Jiancheng Zhao, Yinqiang Zheng

机构 * Graduate School of Information Science and Technology, The University of Tokyo(信息科学与技术研究生院,东京大学)

AI总结 提出统一框架,联合利用全局快门模糊与卷帘快门畸变的互补特性,通过双快门设置和双流运动解释模块实现高速视频重建。

Comments Accepted by TPAMI

URL PDF HTML 收藏
2503.20066 2026-05-25 cs.RO cs.CV

Learning Scene-Level Signed Directional Distance Function with Ellipsoidal Priors and Neural Residuals

学习场景级有符号方向距离函数:结合椭球先验与神经残差

Zhirui Dai, Hojoon Shin, Yulun Tian, Ki Myung Brian Lee, Nikolay Atanasov

机构 * Department of Electrical and Computer Engineering, University of California San Diego(加州大学圣地亚哥分校电气与计算机工程系) Brain Corporation(Brain公司) Robotics Department, University of Michigan(密歇根大学机器人系)

AI总结 提出有符号方向距离函数(SDDF),结合显式椭球先验和隐式神经残差,实现高效准确的场景级几何重建与可微渲染。

Journal ref 2026 IEEE Transactions on Pattern Analysis and Machine Intelligence

URL PDF HTML 收藏
2501.00677 2026-05-22 cs.LG cs.CV cs.IT cs.NA math.IT math.NA stat.ML

Deeply Learned Robust Matrix Completion for Large-scale Low-rank Data Recovery

深度学习鲁棒矩阵补全用于大规模低秩数据恢复

HanQin Cai, Chandra Kundu, Jialin Liu, Wotao Yin

机构 * School of Data, Mathematical, and Statistical Sciences and the Department of Computer Science, University of Central Florida(数据、数学与统计科学学院和计算机科学系,中央佛罗里达大学) School of Data, Mathematical, and Statistical Sciences, University of Central Florida(数据、数学与统计科学学院,中央佛罗里达大学) Damo Academy, Alibaba US(阿里云美国研究院)

AI总结 本文提出了一种可扩展且可学习的非凸方法,即学得鲁棒矩阵补全(LRMC),用于大规模鲁棒矩阵补全问题,该方法具有低计算复杂度和线性收敛性,并通过深度展开有效学习自由参数以实现最优性能,同时在合成数据集和实际应用中验证了其优越的实验性能。

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, 48(6): 6541-6556, 2026

URL PDF HTML 收藏
2605.20963 2026-05-21 cs.CV

Towards UAV Detection in the Real World: A New Multispectral Dataset UAVNet-MS and a New Method

面向现实世界的无人机检测:一个新的多光谱数据集UAVNet-MS和一个新方法

Yihang Luo, Jun Chen, Chao Xiao, Yingqian Wang, Zhaoxu Li, Qiang Ling, Xu He, Nuo Chen, Gaowei Guo, Hongge Li, Miao Li, Longguang Wang, Yulan Guo, Li Liu, Wei An, Zhijie Chen

机构 * College of Electronic Science and Technology, National University of Defense Technology(电子科学与技术学院,国防科技大学) Aviation University of Air Force(空军航空大学) Sun Yat-sen University(中山大学)

AI总结 本文提出了一种新的多光谱数据集UAVNet-MS和一种新的方法MFDNet,用于细粒度小无人机的检测,解决了传统RGB系统在小尺度下的性能问题。

Comments submitted to IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

URL PDF HTML 收藏
2605.19014 2026-05-20 cs.LG econ.EM stat.ML

SAGA: A Sequence-Adaptive Generative Architecture for Multi-Horizon Probabilistic Forecasting with Adaptive Temporal Conformal Prediction

SAGA:一种序列自适应的生成架构,用于多时间跨度概率预测的自适应时间符合预测

Gustav Olaf Yunus Laitinen-Fredriksson Lundström-Imanov, Hafize Gonca Cömert

机构 * Department of Economics, Stockholm University(斯德哥尔摩大学经济系) Institute of Social Sciences, Faculty of Economics and Administrative Sciences, Süleyman Demirel University(苏莱曼·德米雷尔大学社会科学学院,经济学与行政科学学院)

AI总结 本文提出SAGA,一种用于不规则表格面板序列的解码器-only transformer,结合分割符合校准包装器,提供个体层面的预测区间,并保证有限样本边缘覆盖。SAGA在瑞典LISA登记处的纵向数据上训练,预测了1到30年的年度劳动收入,并通过蒙特卡洛方法汇总成现值寿命收入分布。与传统参数过程和表格和循环基线相比,SAGA在10年时间跨度上将连续排名概率分数减少了31.9%,在20年时间跨度上将平均绝对误差减少了37.7%。符合区间在边缘情况下覆盖率为0.4个百分点,在最差的人口子群体中为2.4个百分点。重建的寿命收入基尼系数为0.327,与部分观测的真实值0.341和GKOS估计值0.378相比。模型权重、校准表和合成等价数据集已发布,供在保护的SCB MONA环境中外的复制使用。

Comments 14 pages, 3 figures, 12 tables, 5 appendices, 45 references. Submitted to IEEE TPAMI. Source code at https://github.com/olaflaitinen/saga (archived: doi:10.5281/zenodo.20260366). Synthetic equivalent dataset: doi:10.5281/zenodo.20260287. Empirical work conducted on the Swedish LISA register via SCB MONA (project SCB-MONA-2026-147); ethical approval Swedish Ethical Review Authority 2026-04127-01

URL PDF HTML 收藏
2605.16779 2026-05-19 cs.CV cs.AI

A Holistic Method for Superquadric Fitting Using Unsupervised Clustering Analysis

一种基于无监督聚类分析的超二次曲面拟合整体方法

Mingyang Zhao, Sipu Ruan, Xiaohong Jia

机构 * State Key Laboratory of Mathematical Sciences, Academy of Mathematics and Systems Science, Chinese Academy of Sciences(数学科学国家重点实验室,数学与系统科学学院,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Robotics Institute, School of Mechanical Engineering and Automation, Beihang University(北京航空航天大学机械工程与自动化学院机器人研究所)

AI总结 本文提出了一种新的方法,用于在存在噪声和异常值的情况下对点云进行超二次曲面拟合,通过无监督聚类分析重新定义问题,实现了刚性和变形超二次曲面的一体化拟合,同时提供了闭式解析解和收敛性证明。

Comments 20 pages, Code: https://github.com/zikai1/SuperquadricFitting

Journal ref IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2026

URL PDF HTML 收藏
2510.16046 2026-05-19 physics.soc-ph cs.CY cs.SI

CARDIO-Affect: A Hamiltonian-Variability Framework for Spatio-Temporal Emotional Pattern Recognition with Manifold-Based Individual and Group Profiling

CARDIO-Affect:一种基于哈密顿变异性框架的时空情感模式识别方法,结合基于流形的个体和群体分析

Xiao Sun

AI总结 本文提出CARDIO-Affect框架,通过哈密顿变分理论分析长期情感动态,结合流形学习实现个体和群体情感识别,验证了复杂系统中情感的多稳定性、弱混沌等特征。

Comments v2: Major revision; supersedes v1 ('Neuroticism Paradox', 2025) after FDR-aware re-validation. New complex-systems framework, 6 propositions, three falsifiable paradoxes, Class A AUROC 0.984+/-0.012 matching Granger. Companion: arXiv:2510.15221 (WELD). 23 pages. Submitted to IEEE TPAMI

URL PDF HTML 收藏
2411.17917 2026-05-19 cs.CV cs.RO

DECODE: Domain-aware Continual Domain Expansion for Motion Prediction

DECODE:面向领域的持续领域扩展用于运动预测

Boqi Li, Haojie Zhu, Henry X. Liu

机构 * Department of Civil and Environmental Engineering, University of Michigan(密歇根大学土木与环境工程系)

AI总结 DECODE提出一种持续学习框架,通过预训练模型逐步扩展领域专用模型,结合超网络和流机制实现高效模型选择与不确定性估计,有效降低遗忘率并提升预测精度。

Comments This work has been published in IEEE TPAMI Early Access

URL PDF HTML 收藏
2406.13187 2026-05-19 cs.LG

Decouple then Converge: Handling Unknown Unlabeled Distributions in Long-Tailed Semi-Supervised Learning

解耦然后收敛:处理长尾半监督学习中未知的未标记分布

Kai Gan, Tong Wei, Min-Ling Zhang

机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, China(教育部计算机网络与信息集成重点实验室(东南大学))

AI总结 本文提出DeCon方法,通过解耦学习分支处理长尾半监督学习中未标记数据分布未知的问题,通过两个分支互补提升整体性能。

Comments TPAMI Accepted

URL PDF HTML 收藏
2605.15640 2026-05-18 cs.CV

Learning Disentangled Representations for Generalized Multi-view Clustering

学习解耦表示以实现通用多视图聚类

Xin Zou, Ruimeng Liu, Chang Tang, Zhenglai Li, Xinwang Liu, Kunlun He, Wanqing Li

机构 * AI Thrust, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)人工智能方向) School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院) School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院) School of Computer, National University of Defense Technology(国防科技大学计算机学院) Medical Big Data Research Center, Medical Engineering Laboratory of Chinese PLA General Hospital(中国人民解放军总医院医学大数据研究中心,医学工程实验室) School of Computing and Information Technology, University of Wollongong(沃林根大学计算与信息学院)

AI总结 本文提出GMAE框架,通过解耦表示学习保留多视图互补性,提升聚类效果。实验表明其在完整和不完整多视图聚类任务中均优于现有方法。

Comments accepted by IEEE TPAMI 2026 (IEEE Transactions on Pattern Analysis and Machine Intelligence)

URL PDF HTML 收藏
2605.01852 2026-05-18 cs.CV

DP-SfM: Dual-Pixel Structure-from-Motion without Scale Ambiguity

DP-SfM:无尺度模糊的双像素结构从运动

Lilika Makabe, Kohei Ashida, Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita

机构 * Graduate School of Information Science and Technology, The University of Osaka(信息科学与技术研究生院,大阪大学)

AI总结 本文提出DP-SfM方法,利用双像素传感器的模糊特性自动解决尺度模糊问题,无需参考物体或先验校准,通过深度图与模糊核优化实现绝对尺度估计。

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026

URL PDF HTML 收藏
2605.14750 2026-05-15 cs.CR cs.AI

EVA: Editing for Versatile Alignment against Jailbreaks

EVA:针对对抗性攻击的多功能对齐编辑

Yi Wang, Hongye Qiu, Yue Xu, Sibei Yang, Zhan Qin, Minlie Huang, Wenjie Wang

机构 * ShanghaiTech University(上海科技大学) Sun Yat-sen University(中山大学) State Key Laboratory of Blockchain and Data Security(区块链与数据安全国家重点实验室) Tsinghua University(清华大学)

AI总结 本文提出EVA框架,通过直接模型编辑提升安全对齐,有效中和有害行为,实验显示其在LLMs和VLMs中对抗攻击效果优于现有方法。

Comments IEEE TPAMI 2026

URL PDF HTML 收藏
2605.09020 2026-05-14 cs.CV

The Direct Integration Theorem: A Rigorous Framework for Consistent Discrete Solutions of the Inverse Radon Problem

直接积分定理:逆Radon问题一致离散解的严格框架

Mikhail G. Mozerov

机构 * Institute for Information Transmission Problems, Russian Academy of Sciences(信息传输问题研究所,俄罗斯科学院)

AI总结 本文提出直接积分定理,通过经典中心切片定理推导,为CT中连续到离散转换提供数学一致方法,避免频域插值和常规梯度滤波,解决零频奇点和频域插值误差问题,实现高精度图像重建。

Comments Submitted to IEEE TPAMI. Code and data available at https://github.com/Mozerov-iitp/radon-dit/

URL PDF HTML 收藏
2503.23947 2026-05-13 cs.CV

Spectral-Adaptive Modulation Networks for Visual Perception

用于视觉感知的谱适应调制网络

Guhnoo Yun, Juhan Yoo, Kijung Kim, Jeongho Lee, Paul Hongsuck Seo, Dong Hwan Kim

机构 * Korea University (KU)(韩国大学) Korea Institute of Science and Technology (KIST)(韩国科学技术院) Dong-A University(东洋大学)

AI总结 本文通过图谱分析比较2D卷积与自注意力的频域特性,提出谱适应调制机制,开发出SPANetV2模型,在多个视觉任务中表现优异。

Comments Accepted to IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

URL PDF HTML 收藏
2304.09479 2026-05-13 cs.CV cs.GR cs.LG

DiFaReli++: Diffusion Face Relighting with Consistent Cast Shadows

DiFaReli++: 基于扩散的面部光照重建与一致阴影生成

Puntawat Ponglertnapakorn, Nontawat Tritrong, Supasorn Suwajanakorn

机构 * School of Information Science and Technology, Vidyasirimedhi Institute of Science and Technology(信息科学与技术学院,维达亚西里米迪科学技术研究所)

AI总结 本文提出一种单视图面部光照重建方法,通过条件扩散隐式模型实现无光照地面真值训练,实现真实光照下的阴影一致性。

Comments Published in IEEE TPAMI (vol. 48, no. 5, May 2026). This is an extended version of the ICCV 2023 paper (DiFaReli)

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 48, no. 5, pp. 5068-5082, May 2026

URL PDF HTML 收藏
2605.10717 2026-05-12 cs.LG cs.CV

Heteroscedastic Diffusion for Multi-Agent Trajectory Modeling

异方差扩散用于多智能体轨迹建模

Guillem Capellera, Antonio Rubio, Luis Ferraz, Antonio Agudo

机构 * Institut de Robòtica i Informàtica Industrial, CSIC-UPC(机器人与信息学院,CSIC-UPC)

AI总结 本文提出U2Diffine模型,通过异方差不确定性估计提升多智能体轨迹补全与预测性能,结合Taylor近似和RankNN实现高效推理与误差概率评估,优于现有方法。

Comments Accepted to IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI). Extended version of arXiv:2503.18589 (CVPR 2025)

URL PDF HTML 收藏
2605.08183 2026-05-12 cs.CV cs.LG

Sparsity Hurts: Simple Linear Adapter Can Boost Generalized Category Discovery

稀疏性受损:简单的线性适配器可提升通用类别发现

Bo Ye, Kai Gan, Tong Wei, Min-Ling Zhang

机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, China(计算机网络与信息集成重点实验室(东南大学),教育部,中国)

AI总结 本文提出LAGCD方法,通过在每个ViT块中嵌入残差线性适配器,解决传统方法在灵活性和过拟合问题上的不足,提升通用类别发现性能。

Comments Submitted to IEEE TPAMI

URL PDF HTML 收藏
2605.07378 2026-05-11 cs.LG

Zero-Shot Neural Network Evaluation with Sample-Wise Activation Patterns

基于样本激活模式的零样本神经网络评估

Yameng Peng, Andy Song, HaythamM. Fayek, Vic Ciesielski, Xiaojun Chang

机构 * School of Computing Technologies, RMIT University(计算技术学院,皇家墨尔本理工大学) Department of Electronic Engineering and Information Science, University of Science and Technology of China(电子工程与信息科学系,中国科学技术大学)

AI总结 本文提出SWAP及其衍生指标SWAP-Score,用于评估神经网络性能,克服了传统零样本指标的局限性,展现强预测能力。

Comments Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence. This article is a journal extension of arXiv:2403.04161

URL PDF HTML 收藏
2605.00906 2026-05-05 cs.CV cs.AI cs.LG

Generalized Category Discovery under Domain Shifts: From Vision to Vision-Language Models

在域偏移下进行广义类别发现:从视觉模型到视觉-语言模型

Hongjun Wang, Po Hu, Kai Han

机构 * School of Computing and Data Science, The University of Hong Kong(计算与数据科学学院,香港大学)

AI总结 本文研究在域偏移下的广义类别发现问题,提出三种适应基础模型的框架,从自监督视觉模型到视觉-语言模型,通过多级特征提取和互信息最小化等方法提升性能。

Comments Submission to TPAMI

URL PDF HTML 收藏
2605.00474 2026-05-04 cs.CV

From Local to Global to Mechanistic: An iERF-Centered Unified Framework for Interpreting Vision Models

从局部到全局到机制:一种以iERF为中心的统一框架,用于解释视觉模型

Yearim Kim, Sangyu Han, Nojun Kwak

机构 * Seoul National University(首尔国立大学)

AI总结 本文提出一种以iERF为中心的统一框架,通过点特征向量和实例特定的有效感受野,统一了局部、全局和机制可解释性,提高了解释的准确性和鲁棒性。

Comments Accepted to IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2026

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026

URL PDF HTML 收藏
2604.26917 2026-04-30 cs.CV

AnimateAnyMesh++: A Flexible 4D Foundation Model for High-Fidelity Text-Driven Mesh Animation

AnimateAnyMesh++: 一种灵活的4D基础模型用于高质量文本驱动的网格动画

Zijie Wu, Chaohui Yu, Fan Wang, Xiang Bai

机构 * School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院) DAMO Academy, Alibaba Group(阿里巴巴达摩院) Hupan Lab, Hangzhou, China(湖畔实验室) School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院)

AI总结 本文提出AnimateAnyMesh++,通过扩展数据集、改进架构和生成能力,实现高质量文本驱动的网格动画,提升了轨迹重建和几何保真度。

Comments 14 pages, TPAMI submission, code url: https://github.com/JarrentWu1031/AnimateAnyMesh-pp

URL PDF HTML 收藏
2501.02200 2026-04-22 cs.NE cs.AI cs.CV cs.LG

Learning Evolution via Optimization Knowledge Adaptation

通过优化知识适应学习进化

Chao Wang, Lingling Li, Licheng Jiao, Jiaxuan Zhao, Fang Liu, Shuyuan Yang

机构 * Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education(教育部智能感知与图像理解重点实验室) International Research Center for Intelligent Perception and Computation(智能感知与计算国际研究中心) Xidian University(西安电子科技大学)

AI总结 本文提出OKAEM模型,通过注意力机制参数化进化算子,实现优化知识的预训练和自适应优化,提升进化算法的知识转移与在线适应能力。

Comments This work has been accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence

URL PDF HTML 收藏
2604.18857 2026-04-22 cs.LG cs.CV

Task Switching Without Forgetting via Proximal Decoupling

通过近邻解耦实现无需遗忘的任务切换

Pourya Shamsolmoali, Masoumeh Zareapoor, Eric Granger, William A. P. Smith, Yue Lu

机构 * Department of Computer Science, University of York, UK(英国约克大学计算机科学系) SEIEE, Shanghai Jiao Tong University, China(上海交通大学SEIEE学院) LIVIA, Dept. of Systems Engineering, ETS Montreal, Canada(加拿大蒙特利尔工程学院LIVIA部门) School of Communication and Electronic Eng., East China Normal University, China(华东师范大学通信与电子工程学院)

AI总结 本文提出通过操作符分裂分离任务学习与稳定性维持,提升持续学习的稳定性和适应性,无需使用回放缓冲区、贝叶斯采样或元学习组件。

Comments Submitted to IEEE TPAMI January 2026

URL PDF HTML 收藏
2604.18842 2026-04-22 cs.CV

Multi-Domain Learning with Global Expert Mapping

多域学习与全局专家映射

Pourya Shamsolmoali, Masoumeh Zareapoor, Huiyu Zhou, Oscar Mendez, Dacheng Tao, Xuelong Li

机构 * Department of Computer Science, University of York, UK(英国约克大学计算机科学系) School of Computing and Mathematical Sciences, University of Leicester, UK(英国莱斯特大学计算与数学科学学院) Centre for Vision Speech and Signal Processing, University of Surrey, UK(英国萨里大学视觉语音与信号处理中心) College of Computing and Data Science at Nanyang Technological University, Singapore(新加坡南洋理工大学计算与数据科学学院) Institute of Artificial Intelligence (TeleAI) of China Telecom(中国电信人工智能研究院(TeleAI))

AI总结 本文提出GEM框架,通过全局调度器替代传统路由器,解决多域学习中专家分配不均问题,提升模型在罕见域和分布外域的表现。

Comments Submitted to IEEE TPAMI on August 2025

URL PDF HTML 收藏
2604.14726 2026-04-22 cs.LG cs.AI

Catching Every Ripple: Enhanced Anomaly Awareness via Dynamic Concept Adaptation

捕捉每一个涟漪:通过动态概念适应提升异常意识

Jiaqi Zhu, Shaofeng Cai, Jie Chen, Fang Deng, Beng Chin Ooi, Wenqiao Zhang

机构 * School of Automation and the National Key Laboratory of Autonomous Intelligent Unmanned Systems, Beijing Institute of Technology(自动化学院和自主智能无人系统国家重点实验室,北京理工大学) School of Computing, National University of Singapore(computing 学院,新加坡国立大学) Harbin Institute of Technology(哈尔滨工业大学) Digital Media Computing & Design Lab, Zhejiang University(数字媒体计算与设计实验室,浙江大学)

AI总结 本文提出DyMETER框架,通过动态概念适应机制提升在线异常检测的适应性与效率,采用超网络生成实例感知参数位移,并引入轻量级进化控制器和动态阈值优化模块以实现鲁棒和可解释的适应。

Comments Accepted by IEEE TPAMI

URL PDF HTML 收藏
2502.16161 2026-04-22 cs.CV cs.CL

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models

OmniParser V2:结构化思维点用于统一的视觉文本解析及其在多模态大语言模型中的通用性

Wenwen Yu, Zhibo Yang, Jianqiang Wan, Sibo Song, Jun Tang, Wenqing Cheng, Yuliang Liu, Xiang Bai

机构 * School of Information Science and Engineering, East China University of Science and Technology(东华大学信息科学与工程学院) School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院) School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院) School of Electronic Information and Communications, Huazhong University of Science and Technology(华中科技大学电子信息与通信学院) Alibaba Group(阿里巴巴集团)

AI总结 本文提出OmniParser V2,通过结构化思维点提示方案统一视觉文本解析任务,简化流程并提升性能,在多个数据集上取得最佳结果,并验证其在多模态大语言模型中的通用性。

Comments Accepted by IEEE TPAMI

URL PDF HTML 收藏
2604.15654 2026-04-20 cs.CV

From Zero to Detail: A Progressive Spectral Decoupling Paradigm for UHD Image Restoration with New Benchmark

从零到细节:一种渐进式频谱解耦范式用于超高清图像修复的新基准

Chen Zhao, Yunzhe Xu, Zhizhou Chen, Enxuan Gu, Kai Zhang, Xiaoming Liu, Jian Yang, Ying Tai

机构 * State Key Laboratory of Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学) Department of Computer Science and Engineering, Michigan State University(计算机科学与工程系,密歇根州立大学) School of Computer Science and Technology, Dalian University of Technology(计算机科学与技术学院,大连理工大学)

AI总结 本文提出渐进式频谱解耦范式,通过零频增强、低频修复和高频细化三个阶段提升超高清图像修复效果,并构建了高质量基准数据集LSUHDIR。

Comments TPAMI

URL PDF HTML 收藏
2501.05281 2026-04-20 cs.CV cs.LG

Comparison Study: Glacier Calving Front Delineation in Synthetic Aperture Radar Images With Deep Learning

比较研究:利用深度学习在合成孔径雷达图像中进行冰川崩解前端 delineation

Nora Gourmelon, Konrad Heidler, Erik Loebel, Daniel Cheng, Julian Klink, Anda Dong, Fei Wu, Noah Maul, Moritz Koch, Marcel Dreier, Dakota Pyles, Thorsten Seehaus, Matthias Braun, Andreas Maier, Vincent Christlein

机构 * Department of Computer Science, Friedrich-Alexander-Universität Erlangen-Nürnberg(弗里德里希-亚历山大-埃朗根-纽伦堡大学计算机科学系) School of Engineering and Design, Technische Universität München(慕尼黑技术大学工程与设计学院) Institut für Planetare Geodäsie, Technische Universität Dresden(德累斯顿技术大学行星大地测量研究所) Jet Propulsion Laboratory, California Institute of Technology(加州理工学院喷气推进实验室) Institut für Geographie, Friedrich-Alexander-Universität Erlangen-Nürnberg(弗里德里希-亚历山大-埃朗根-纽伦堡大学地理研究所)

AI总结 本研究比较了深度学习系统在合成孔径雷达图像中进行冰川崩解前端 delineation 的性能,发现深度学习系统存在高达221米的误差,而人工标注者误差仅为38米,凸显了进一步研究的必要性。

Comments Accepted as short paper in IEEE Transactions on Pattern Analysis and Machine Intelligence

URL PDF HTML 收藏
2604.14795 2026-04-17 cs.RO

Keep It CALM: Toward Calibration-Free Kilometer-Level SLAM with Visual Geometry Foundation Models via an Assistant Eye

保持冷静:通过视觉几何基础模型实现无校准的千米级SLAM

Tianjun Zhang, Fengyi Zhang, Tianchen Deng, Lin Zhang, Hesheng Wang

机构 * School of Automation and Intelligent Sensing, and Shanghai Key Laboratory of Navigation and Location Based Services, Shanghai Jiao Tong University(自动化与智能感知学院,导航与位置服务上海市重点实验室,上海交通大学) School of Electrical Engineering and Computer Science, The University of Queensland(电气工程与计算机科学学院,昆士兰大学) School of Computer Science and Technology, Tongji University(计算机科学与技术学院,同济大学)

AI总结 本文提出CAL2M框架,通过引入'助手眼'消除尺度模糊,利用epipolar引导模型修正内参和姿态,结合锚点传播实现全局一致的映射,解决传统方法在千米级SLAM中的精度问题。

Comments 19 pages, 8 figures, submitted to IEEE TPAMI

URL PDF HTML 收藏