arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision

2026-08-11 至 2026-08-11 收录 8
2608.08867 2026-08-11 cs.CV 新提交

Zero-Shot Traffic Accident Detection via a Coarse-to-Fine VLM-Tracking Pipeline

基于粗到细VLM跟踪流水线的零样本交通事故检测

Dipit Saha, Shah Mohammad Abdul Mannan, Mohammad Raihan Rashid, Ruwad Naswan, Ahnaf Tahmid

机构 * Bangladesh University of Engineering and Technology(孟加拉工程技术大学)

AI总结 本文提出一种无需训练的双通粗到细流水线,结合Qwen3-VL-32B-Instruct、YOLO11x与BoT-SORT,在零样本约束下于ACCIDENT @ CVPR基准测试中实现22%相对优势,达成0.504的三方调和均值得分。

Comments Accepted at the AUTOPILOT Workshop, CVPR 2026, Denver, CO

URL PDF HTML 收藏
2602.20537 2026-08-11 cs.CV

PFGNet: A Fully Convolutional Frequency-Guided Peripheral Gating Network for Efficient Spatiotemporal Predictive Learning

PFGNet: 一种全卷积频率引导边缘门控网络,用于高效的时空预测学习

Xinyong Cai, Changbin Sun, Yong Wang, Hongyu Yang, Yuankai Wu

机构 * College of Computer Science, Sichuan University(四川大学计算机科学学院) Department of Data and Systems Engineering, The University of Hong Kong(香港大学数据与系统工程系)

AI总结 PFGNet通过像素级频率引导门控动态调节感受野,实现高效的时空预测学习,无需递归或注意力机制,在多个数据集上取得SOTA性能。

Comments Accepted to CVPR 2026

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026, pp. 38848-38858

URL PDF HTML 收藏
2604.15941 2026-08-11 cs.CV cs.GR

Neural Gabor Splatting: Enhanced Gaussian Splatting with Neural Gabor for High-frequency Surface Reconstruction

神经戈尔巴式点云:通过神经戈尔巴增强高频率表面重建

Haato Watanabe, Nobuyuki Umetani

机构 * The University of Tokyo(东京大学)

AI总结 本文提出神经戈尔巴点云方法,通过在高斯点云中加入轻量级多层感知机来增强颜色变化的表示,减少点云数量,实现高频率表面的精确重建。

Comments Accepted to CVPR 2026

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4932-4941, 2026

URL PDF HTML 收藏
2510.20512 2026-08-11 cs.CV 版本更新

Adversarial Concept Distillation for One-Step Diffusion Personalization

对抗概念蒸馏用于一步扩散个性化

Yixiong Yang, Tao Wu, Senmao Li, Shiqi Yang, Yaxing Wang, Joost van de Weijer, Kai Wang

机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) Computer Vision Center(计算机视觉中心) Universitat Autònoma de Barcelona(巴塞罗那自治大学) VCIP, CS, Nankai University(南开大学计算机学院视觉计算与智能感知实验室) Program of Computer Science, City University of Hong Kong (Dongguan)(香港城市大学(东莞)计算机科学项目) City University of Hong Kong(香港城市大学)

AI总结 本文提出OPAD框架,结合教师-学生蒸馏与对抗监督,通过多步扩散模型作为教师,一步学生模型联合训练,以实现一步扩散模型的高效个性化。

Comments Accepted to CVPR 2026 Findings

URL PDF HTML 收藏
2603.03920 2026-08-11 cs.LG cs.AI

BD-Merging: Bias-Aware Dynamic Model Merging with Evidence-Guided Contrastive Learning

BD-Merging: 偏差感知动态模型合并与证据引导对比学习

Yuhan Xie, Chen Lyu

机构 * MoE Key Laboratory of Interdisciplinary Research of Computation and Economics(计算与经济学交叉学科联合实验室)

AI总结 BD-Merging通过引入证据引导对比学习,解决多任务学习中分布偏移导致的预测偏差问题,提升模型鲁棒性和泛化能力。

Comments Accepted by CVPR 2026

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026

URL PDF HTML 收藏
2411.16758 2026-08-11 cs.CV 版本更新

Motion-Aware Animatable Gaussian Avatars Deblurring

具有运动感知的可动画高斯人像去模糊

Muyao Niu, Yifan Zhan, Qingtian Zhu, Zhuoxiao Li, Wei Wang, Zhihang Zhong, Xiao Sun, Yinqiang Zheng

机构 * The University of Tokyo(东京大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学)

AI总结 本文提出了一种从模糊视频中直接重建清晰3D人体高斯人像的方法,结合了基于物理的模糊模型和运动模型,以解决运动引起的模糊问题。

Comments Accepted at CVPR 2026, Codes: https://github.com/MyNiuuu/MAD-Avatar

URL PDF HTML 收藏
2603.02270 2026-08-11 cs.CV 版本更新

From Visual to Multimodal: Systematic Ablation of Encoders and Fusion Strategies in Animal Identification

从视觉到多模态:动物识别中编码器和融合策略的系统消融

Vasiliy Kudryavtsev, Kirill Borodin, German Berezin, Kirill Bubenchikov, Grach Mkrtchian, Alexander Ryzhkov

机构 * Faculty of IT, Technical University of Communication(信息科技学院,通信技术大学) AI lab, Avito(人工智能实验室,Avito)

AI总结 本研究提出多模态验证框架,通过合成文本描述提升视觉特征,利用门控融合机制在动物识别中实现84.28%的准确率,比单一模态基线提升11%。

Comments Accepted to the FGVC13 Workshop at CVPR 2026. And published at MDPI Journal of Imaging (see at https://www.mdpi.com/2313-433X/12/1/30)

Journal ref Journal of Imaging (2026) 12, no. 1: 30

URL PDF HTML 收藏
2506.07223 2026-08-11 cs.AI 版本更新

Reflex First, Reflect Later: Latency-Aware Embodied LLM Agents for Dynamic Response

先反射,后反思:面向动态响应的延迟感知具身大语言模型智能体

Yangqing Zheng, Shunqi Mao, Dingxin Zhang, Weidong Cai

机构 * School of Computer Science, The University of Sydney(计算机科学学院,悉尼大学)

AI总结 该研究针对动态环境中具身LLM智能体的推理延迟问题,提出RRARA智能体及相关评估指标,通过时间转换机制与预规划器实现决策质量与响应能力的平衡。

Comments Accepted by the CVPR 2025 Embodied AI Workshop

URL PDF HTML 收藏