arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

International Conference on Computer Vision · 会议 · Computer Vision

共收录 4779 篇
2609.37870 2026-09-30 cs.CV 新提交

Learning from synthetic photorealistic raindrop for single image raindrop removal

学习合成逼真雨滴用于单图像雨滴去除

Zhixiang Hao, Shaodi You, Yu Li, Kunming Li, Feng Lu

机构 * State Key Laboratory of VR Technology and Systems, Beihang University(北京航空航天大学虚拟现实技术与系统国家重点实验室) ; Data61-CSIRO(澳大利亚联邦科学与工业研究组织Data61) ; Tencent(腾讯) ; Australian National University(澳大利亚国立大学) ; Peng Cheng Laboratory(鹏城实验室)

AI总结 针对雨天图像中雨滴干扰问题,提出首个基于物理渲染的合成逼真雨滴数据集,并设计感知折射与模糊的检测网络及恢复结构的去除网络,实现单图像雨滴去除的最先进性能。

Comments 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06968 2026-09-29 cs.CV cs.GR

3D Gaussian Splatting with Fisheye Images: Field of View Analysis and Depth-Based Initialization

基于鱼眼图像的3D高斯点撒:视场分析与基于深度的初始化

Ulas Gunes, Matias Turkulainen, Mikhail Silaev, Juho Kannala, Esa Rahtu

机构 * Department of Computing Sciences, Tampere University(塔尔基大学计算机科学系) ; Department of Computer Science, Aalto University(阿尔托大学计算机科学系) ; Department of Computer Science and Engineering, University of Oulu(奥卢大学计算机科学与工程系)

AI总结 本文首次评估了基于鱼眼图像的3D高斯点撒方法,通过对比不同视场下的重建效果,发现160度视场表现最佳,并引入基于深度的UniK3D方法提升宽角重建性能。

Comments VISAPP 2026 Accepted Camera Ready Version

Journal ref Proceedings of the 21st International Conference on Computer Vision Theory and Applications (VISAPP), Vol. 3, pp. 229-236, SciTePress, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12906 2026-09-29 cs.CV 版本更新

CATSplat: Context-Aware Transformer with Spatial Guidance for Generalizable 3D Gaussian Splatting from A Single-View Image

CATSplat:用于从单视图图像进行可泛化3D高斯泼溅的带空间引导的上下文感知Transformer

Wonseok Roh, Hwanhee Jung, Jong Wook Kim, Seunggwan Lee, Innfarn Yoo, Andreas Lugmayr, Seunggeun Chi, Karthik Ramani, Sangpil Kim

机构 * Korea University(高丽大学) ; Google(谷歌公司) ; Purdue University(普渡大学)

AI总结 提出CATSplat框架,通过引入视觉语言模型的文本指导和3D点特征的空间指导,突破单目设置限制,实现高质量的单视图3D场景重建与新视图合成。

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07990 2026-09-18 cs.CV

GraphEnet: Event-driven Human Pose Estimation with a Graph Neural Network

GraphEnet: 基于图神经网络的事件驱动人体姿态估计

Gaurvi Goyal, Pham Cong Thuong, Arren Glover, Masayoshi Mizuno, Chiara Bartolozzi

机构 * Maastricht University(马斯特里赫特大学) ; Istituto Italiano di Tecnologia(意大利技术研究院) ; Sony Interactive Entertainment Inc.(索尼互动娱乐公司)

AI总结 针对事件相机的低延迟低能耗特性及人体姿态估计需求,提出基于图神经网络的GraphEnet,采用线条中间表示、偏移向量学习与置信度池化实现单人2D姿态高频估计,为该领域首项GNN应用研究且代码开源。

Journal ref 2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.18490 2026-09-17 cs.CV 新提交

Learning A Unified Template for Gait Recognition

学习统一模板用于步态识别

Panjian Huang, Saihui Hou, Junzhou Huang, Yongzhen Huang

机构 * Beijing Normal University(北京师范大学) ; The University of Texas at Arlington(德克萨斯大学阿灵顿分校)

AI总结 针对步态识别中语义不一致和均匀性问题,提出基于扩散模型思想的Origins框架,通过统一模板的生成与表示联合学习,在多个基准数据集上取得优越性能。

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.16306 2026-09-16 cs.CV cs.LG 新提交

Sequence Recognition in Bharatnatyam dance

Bharatnatyam舞蹈中的序列识别

Himadri Bhuyan, Rohit Dhaipule, Partha Pratim Das

机构 * Indian Institute of Technology Kharagpur(印度理工学院卡拉格普尔分校)

AI总结 本文提出一种基于CNN和SVM识别Bharatnatyam舞蹈中Adavu序列的方法,利用编辑距离匹配序列,准确率达98%,并测试了所有变体的可扩展性。

Comments Accepted at 7th International Conference on Computer Vision and Image Processing (CVIP), 2022

Journal ref Computer Vision and Image Processing. CVIP 2022. Communications in Computer and Information Science, vol 1778. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12441 2026-09-16 cs.CV cs.LG

Describe Anything Model for Visual Question Answering on Text-rich Images

面向富文本图像视觉问答的Describe Anything Model

Yen-Linh Vu, Dinh-Thang Duong, Truong-Binh Duong, Anh-Khoi Nguyen, Thanh-Huy Nguyen, Le Thien Phuc Nguyen, Jianhua Xing, Xingjian Li, Tianyang Wang, Ulas Bagci, Min Xu

机构 * AI VIETNAM Lab(AI越南实验室) ; Carnegie Mellon University(卡内基梅隆大学) ; University of Wisconsin - Madison(威斯康星大学麦迪逊分校) ; University of Pittsburgh(匹兹堡大学) ; University of Alabama at Birmingham(阿拉巴马大学伯明翰分校) ; Northwestern University(西北大学)

AI总结 本研究提出DAM-QA框架,利用Describe Anything Model的区域感知能力,通过多区域视图答案聚合机制解决富文本图像视觉问答问题,在6个基准上优于基线,DocVQA提升超7个点,参数更少且性能领先。

Comments 11 pages, 5 figures. Accepted to VisionDocs @ ICCV 2025

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16076 2026-09-11 cs.HC cs.CV 版本更新

Prompting with Sign Parameters for Low-resource Sign Language Instruction Generation

面向低资源手语指令生成的符号参数提示

Md Tariquzzaman, Md Farhan Ishmam, Saiyma Sittul Muna, Md Kamrul Hasan, Hasan Mahmud

机构 * Department of Computer Science and Engineering(计算机科学与工程系)

AI总结 针对低资源手语指令生成,提出首个孟加拉语数据集BdSLIG及符号参数注入提示方法,提升零样本性能,促进包容性。

Comments Accepted at the ICCV 2025 Workshop on Vision Foundation Models and Generative AI for Accessibility (CV4A11y). OpenReview: https://openreview.net/pdf?id=KkVMBkjbra

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22076 2026-09-10 cs.LG 版本更新

Test-time Prompt Refinement for Text-to-Image Models

文本到图像模型的测试时提示词优化

Mohammad Abdul Hafeez Khan, Yash Jain, Siddhartha Bhattacharyya, Vibhav Vineet

机构 * Florida Institute of Technology(佛罗里达理工学院) ; Microsoft Research(微软研究院)

AI总结 提出TIR框架,利用多模态大语言模型在测试时迭代优化提示词,无需训练T2I模型,提升生成图像与提示词的对齐性和视觉连贯性。

Comments Accepted to ICCV 2025, MARS2 Workshop. Total 14 pages, 12 figures and 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16015 2026-09-10 cs.CV

Is Tracking really more challenging in First Person Egocentric Vision?

第一人称自我中心视觉中的跟踪真的更具挑战性吗?

Matteo Dunnhofer, Zaira Manigrasso, Christian Micheloni

机构 * University of Udine(乌迪内大学) ; York University(约克大学)

AI总结 针对现有研究认为第一人称自我中心视觉跟踪更具挑战但评估场景差异大的问题,该研究提出新型基准测试以分离视角与人-物活动领域的影响,明确难点来源以推动任务发展。

Comments 2025 IEEE/CVF International Conference on Computer Vision (ICCV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09563 2026-09-09 cs.CV 版本更新

Robust Optical Flow Computation: A Higher-Order Differential Approach

鲁棒光流计算:一种高阶微分方法

Chanuka Algama, Kasun Amarasinghe

机构 * University of Kelaniya(凯拉尼亚大学) ; Carnegie Mellon University(卡内基梅隆大学)

AI总结 针对大非线性运动下光流估计难题,提出基于二阶泰勒近似的微分方法,在KITTI和Middlebury基准上显著降低平均端点误差。

Comments 8 pages. Revised to match the published VISAPP 2026 version; redundant material was removed

Journal ref Proceedings of the 21st International Conference on Computer Vision Theory and Applications (VISAPP), Volume 3, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29227 2026-09-01 cs.CV 新提交

Using Channel Representations in Regularization Terms: A Case Study on Image Diffusion

在正则化项中使用通道表示:图像扩散的案例研究

Christian Heinemann, Freddie Åström, George Baravdish, Kai Krajsek, Michael Felsberg, Hanno Scharr

机构 * Forschungszentrum Jülich(于利希研究中心) ; Linköping University(林雪平大学)

AI总结 本研究提出基于图像通道表示的新型非线性扩散滤波方法,构建含通道表示权重项的能量泛函,在含混合噪声的图像重建与去噪任务中表现具竞争力。

Journal ref Proceedings of the 9th International Conference on Computer Vision Theory and Applications (VISAPP 2014), vol. 2, pp. 48-55, SciTePress, 2014

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18898 2026-08-18 cs.CV cs.GR 版本更新

GestureLSM: Latent Shortcut based Co-Speech Gesture Generation with Spatial-Temporal Modeling

GestureLSM:基于隐式捷径的时空建模语音同步手势生成

Pinxin Liu, Luchuan Song, Junhua Huang, Haiyang Liu, Junfan Zhu, Chenliang Xu

机构 * University of Rochester(罗切斯特大学) ; University of Tokyo(东京大学)

AI总结 GestureLSM通过时空建模与改进的流匹配方法,在BEAT2数据集上实现了最优语音同步手势生成,同时大幅缩短推理时间,可用于增强数字人和具身智能体。

Comments Accepted to ICCV 2025. Project Page: https://andypinxinliu.github.io/GestureLSM

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16563 2026-08-18 cs.CV 版本更新

SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis

SemTalk:具有帧级语义强调的整体同步言语动作生成

Xiangyue Zhang, Jianfang Li, Jiaxu Zhang, Ziqiang Dang, Jianqiang Ren, Liefeng Bo, Zhigang Tu

机构 * Wuhan University(武汉大学) ; Alibaba(阿里巴巴) ; Zhejiang University(浙江大学)

AI总结 SemTalk通过分离学习基础与稀疏动作并自适应融合,结合coarse2fine交叉注意力等技术,在两个公开数据集上实现了优于现有最优方法的高质量同步言语动作生成。

Comments 11 pages, 8 figures. Accepted to ICCV 2025. Project page: https://xiangyuezhang.com/SemTalk/; code and pretrained models: https://github.com/Xiangyue-Zhang/SemTalk

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025, pp. 13761-13771

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.03725 2026-08-17 cs.CV

I2UV-HandNet: Image-to-UV Prediction Network for Accurate and High-fidelity 3D Hand Mesh Modeling

I2UV-HandNet:用于精准高保真3D手部网格建模的图像到UV预测网络

Ping Chen, Yujin Chen, Dong Yang, Fangyin Wu, Qin Li, Qingpei Xia, Yong Tan

机构 * IQIYI Inc.(爱奇艺公司) ; Wuhan University(武汉大学) ; Technical University of Munich(慕尼黑工业大学)

AI总结 针对现有手部3D重建精度与保真度不足的问题,提出I2UV-HandNet模型,采用首个基于UV的手部形状表示,结合AffineNet与SRNet实现精准高保真3D手部网格建模,在多个基准上达SOTA性能。

Comments Accepted by ICCV2021; 14 pages, 8 figures

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 12929-12938

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03243 2026-08-06 cs.CV 版本更新

MVTOP: Multi-View Transformer-based Object Pose-Estimation

MVTOP:基于多视图的Transformer物体姿态估计

Lukas Ranftl, Felix Brendel, Bertram Drost, Carsten Steger

机构 * MVTec Software GmbH(MVTec软件公司) ; Technical University of Munich(慕尼黑技术大学)

AI总结 MVTOP通过多视图早期融合解决单视图无法解决的姿态歧义问题,利用视线线模型实现多视图几何建模,优于现有单视图和多视图方法。

Comments 9 pages, 7 figures, Accepted as Conference paper to VISAPP 2026

Journal ref Proceedings of the 21st International Conference on Computer Vision Theory and Applications, VISAPP 2026, Marbella, Spain, March 9-11, 2026, Volume 3. SCITEPRESS 2026, ISBN 978-989-758-804-4, pages 181-194

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14171 2026-08-05 cs.LG cs.AI 版本更新

IPPRO: Importance-based Pruning with PRojective Offset for Magnitude-indifferent Structural Pruning

IPPRO:面向幅值无关结构化剪枝的基于重要性的投影偏移剪枝

Jaeheun Jung, Jaehyuk Lee, Yeajin Lee, Donghun Lee

AI总结 针对现有结构化剪枝依赖滤波器幅值的尺度不变性缺陷,提出基于投影几何的IPPRO框架,定义PROscore并关联$L_0$松弛,在多类模型上均优于现有方法,为神经网络压缩提供鲁棒架构无关范式。

Comments ICCV 2025 workshop U&ME 2025 (2nd Workshop and Challenge on Unlearning and Model Editing)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07754 2026-08-05 cs.LG cs.AI 版本更新

One-Point Contraction: Erasing Representational Separability toward Irreversible Deep Forgetting

单点收缩:擦除表示可分性以实现不可逆的深度遗忘

Jaeheun Jung, Bosung Jung, Suhyun Bae, Donghun Lee

机构 * Korea University(韩国大学)

AI总结 本研究发现现有机器遗忘方法可被FM-recovery逆转,提出OPC方法,可实现行为与表示级遗忘,抵御攻击且不牺牲准确率。

Comments ICCV 2025 workshop U&ME 2025 (2nd Workshop and Challenge on Unlearning and Model Editing)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02029 2026-07-30 cs.CV cs.AI

Fake & Square: Training Self-Supervised Vision Transformers with Synthetic Data and Synthetic Hard Negatives

Fake & Square:利用合成数据与合成难负例训练自监督视觉Transformer

Nikos Giakoumoglou, Andreas Floros, Kleanthis Marios Papadopoulos, Tania Stathaki

机构 * Imperial College London(伦敦帝国学院)

AI总结 该研究提出Syn2Co框架,结合合成数据增强与表示空间合成难负例生成两种策略,在DeiT-S和Swin-T架构上验证了合成增强训练对提升视觉表示鲁棒性与迁移性的作用,明确了其前景与局限。

Comments ICCV 2025 Workshop LIMIT

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26043 2026-07-29 cs.LG 新提交

Re-thinking Mammography Transfer Learning: The Dataset-Informed Transfer Learning (DITL) Framework for Breast Cancer Screening and Lesion Diagnosis

重新思考乳腺X线摄影转移学习:用于乳腺癌筛查和病变诊断的数据集知情转移学习(DITL)框架

Adarsh Bhandary Panambur, Siming Bayer, Andreas Maier

机构 * Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg(模式识别实验室,埃尔朗根-纽伦堡弗里德里希-亚历山大大学) ; Siemens Healthineers(西门子医疗)

AI总结 研究针对乳腺X线摄影分类性能提升难题,提出DITL框架,整合数据集难度信号与邻域监督,引入自适应组件,无需超参数调整,在大规模和小数据集上均有出色表现,建立了通用的乳腺X线摄影分类框架。

Comments 16 pages, 1 figure, 5 tables. Accepted and presented at the 10th International Conference on Computer Vision & Image Processing (CVIP 2025), IIT Ropar, India, 10-13 December 2025. The paper is currently in press for inclusion in the official conference proceedings. This preprint corresponds to the submitted manuscript and is made available pending publication of the final proceedings version

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21600 2026-07-28 cs.CV

Locally Controlled Face Aging with Latent Diffusion Models

局部控制的面部老化与潜在扩散模型

Lais Isabelle Alves dos Santos, Julien Despois, Thibaut Chauffier, Sileye O. Ba, Giovanni Palma

机构 * L’Oréal AI Research(欧莱雅人工智能研究所)

AI总结 本文提出利用潜在扩散模型实现局部控制的面部老化,通过细粒度控制提升生成结果的真实性和可控性。

Journal ref 2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2025, pp. 6991-6999

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.16926 2026-07-21 cs.CV 新提交

Splat-based 3D Scene Reconstruction with Extreme Motion-blur

基于体素的极端运动模糊3D场景重建

Hyeonjoong Jang, Dongyoung Choi, Donggun Kim, Woohyun Kang, Min H. Kim

机构 * KAIST(韩国科学技术院) ; HYPERGRAM(超图)

AI总结 研究针对低光照下RGB-D输入的极端运动模糊问题,提出结合高斯体素框架的相机姿态估计与图像去模糊方法,通过对齐帧、调整高斯位置等步骤,提升3D重建质量,优于现有方法,有广泛应用意义。

Journal ref Proceedings of IEEE/CVF International Conference on Computer Vision (ICCV) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18242 2026-07-17 cs.CV 版本更新

GSVisLoc: Generalizable Visual Localization for Gaussian Splatting Scene Representations

GSVisLoc:用于高斯喷溅场景表示的通用视觉定位

Fadi Khatib, Dror Moran, Guy Trostianetsky, Yoni Kasten, Meirav Galun, Ronen Basri

机构 * Weizmann Institute of Science(魏茨曼科学研究所) ; NVIDIA(英伟达)

AI总结 GSVisLoc是用于3D高斯喷溅场景表示的视觉定位方法,通过匹配3D高斯生成的场景特征与图像特征来估计相机位姿,分三步进行,无需修改、再训练或额外图像,在标准基准上性能优且能有效推广到新场景。

Comments Accepted to ICCV 2025 Workshops (CALIPOSE). Project page: https://gsvisloc.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19129 2026-07-14 cs.CV 版本更新

KAMERA: Enhancing Aerial Surveys of Ice-associated Seals in Arctic Environments

KAMERA:增强北极环境中与冰相关海豹的航空调查

Adam Romlein, Benjamin X. Hou, Yuval Boss, Cynthia L. Christman, Stacie Koslovsky, Erin E. Moreland, Jason Parham, Anthony Hoogs

机构 * NOAA NMFS AFSC MML(国家海洋管理局NMFS AFSC MML) ; CICOES(国际极地科学组织) ; University of Washington(华盛顿大学)

AI总结 KAMERA系统用于北极环境中与冰相关海豹的航空调查,通过多相机多光谱同步及实时检测,减少数据集处理时间,能利用多光谱检测目标,数据带元数据,图像和检测结果可映射到世界平面,软件等完全开源。

Comments 10 pages, 9 figures, 4 tables. Code: https://github.com/Kitware/kamera

Journal ref 2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2025, pp. 2183-2192

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13184 2026-07-14 cs.CV

Triad: Empowering LMM-based Anomaly Detection with Vision Expert-guided Visual Tokenizer and Manufacturing Process

Triad: 通过视觉专家引导的视觉标记器和制造过程增强基于LMM的异常检测

Yuanze Li, Shihao Yuan, Haolin Wang, Qizhang Li, Ming Liu, Chen Xu, Guangming Shi, Wangmeng Zuo

机构 * Harbin Institute of Technology(哈尔滨工业大学) ; Pengcheng Lab(鹏城实验室)

AI总结 Triad通过引入视觉专家引导的标记器和制造过程,提升基于LMM的工业异常检测性能。

Journal ref In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 21917-21926. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08875 2026-07-03 cs.AI 版本更新

Causal Explanations for Image Classifiers

图像分类器的因果解释

Hana Chockler, David A. Kelly, Daniel Kroening, Youcheng Sun

机构 * King's College London(伦敦国王学院) ; Amazon.com, Inc.(亚马逊公司) ; Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) ; University of Manchester(曼彻斯特大学)

AI总结 本文提出基于实际因果理论的新型黑盒方法,证明算法终止性并展示其在解释效率和精度上的优势。

Comments Accepted to Journal of Artificial Intelligence Research (JAIR). A subset of the contribution was published in ICCV 2021

Journal ref Journal of Artificial Intelligence Research, vol 86, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08237 2026-07-01 cs.MM cs.AI cs.CV cs.SD eess.AS 版本更新

VGGSounder: Audio-Visual Evaluations for Foundation Models

VGGSounder:基础模型的音视频评估

Daniil Zverev, Thaddäus Wiedemer, Ameya Prabhu, Matthias Bethge, Wieland Brendel, A. Sophia Koepke

机构 * Technical University of Munich, MCML(慕尼黑技术大学,MCML) ; University of Tübingen(图宾根大学) ; Tübingen AI Center(图宾根人工智能中心) ; MPI for Intelligent Systems, ELLIS Institute(智能系统Max Planck研究所,ELLIS研究所)

AI总结 针对VGGSound数据集在音视频基础模型评估中的标签不完整、类别重叠和模态错位等问题,提出重新标注的多标签测试集VGGSounder,并引入模态混淆指标分析模型性能退化。

Comments Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11061 2026-07-01 cs.CV cs.AI 版本更新

Robust 3D-Masked Part-level Editing in 3D Gaussian Splatting with Regularized Score Distillation Sampling

基于正则化分数蒸馏采样的鲁棒3D掩膜部分级编辑在3D高斯泼溅中的应用

Hayeon Kim, Ji Ha Jang, Se Young Chun

AI总结 提出RoMaP框架,通过3D几何感知标签预测生成鲁棒3D掩膜,并引入正则化SDS损失(含SLaMP编辑的L1锚点损失),实现精确且大幅度的局部3D高斯编辑。

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15864 2026-06-30 cs.LG

Improving Rectified Flow with Boundary Conditions

通过边界条件改进校正流

Xixi Hu, Runlong Liao, Keyang Xu, Bo Liu, Yeqing Li, Eugene Ie, Hongliang Fei, Qiang Liu

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校) ; Google(谷歌)

AI总结 本文提出边界约束校正流模型,通过强制边界条件提升生成模型性能,在ImageNet上使用ODE和SDE采样分别提升FID分数8.01%和8.98%。

Comments ICCV 2025

Journal ref Proc. IEEE/CVF International Conference on Computer Vision (ICCV), 2025, pp. 18177-18186

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24847 2026-06-24 cs.CV 新提交

Spherical-to-ERP Epipolar Rectification for Single-Axis Disparity in 360 Stereo

360度立体中单轴视差的球面到ERP极线校正

Sahereh Obeidavi, Dieter Landes

机构 * Faculty of Electrical Engineering and Computer Science, Coburg University of Applied Science(电气工程与计算机科学学院,科堡应用科学大学)

AI总结 针对球面立体图像中极线对应呈曲线导致二维位移的问题,采用球面到等距柱状投影预处理将极线拉直恢复单轴视差,结合RAFT+EACS框架实现实时高精度视差估计。

Comments 7 Pages, 4 Figures, Conference

Journal ref International Conference on Computer Vision and Artificial Intelligence (ICCVAI - 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
↑