arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Winter Conference on Applications of Computer Vision · 会议 · Computer Vision

2025-12-12 至 2025-12-12 共收录 10
2512.10939 2025-12-12 cs.CV

GaussianHeadTalk: Wobble-Free 3D Talking Heads with Audio Driven Gaussian Splatting

GaussianHeadTalk: 无抖动的3D说话头:基于音频驱动的高斯点云法

Madhav Agarwal, Mingtian Zhang, Laura Sevilla-Lara, Steven McDonagh

机构 * University of Edinburgh(爱丁堡大学) University College London(伦敦大学学院)

AI总结 本文提出基于音频驱动的3D说话头生成方法,利用高斯点云法和Transformer模型实现高保真、实时的虚拟角色生成。

Comments IEEE/CVF Winter Conference on Applications of Computer Vision 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10321 2025-12-12 cs.CV

Point2Pose: A Generative Framework for 3D Human Pose Estimation with Multi-View Point Cloud Dataset

Point2Pose:基于多视角点云数据集的3D人体姿态估计生成框架

Hyunsoo Lee, Daeum Jeon, Hyeokjae Oh

机构 * ECE, Seoul National University(首尔国立大学电子与计算机工程系) CS, KAIST(韩国科学技术院计算机科学系) Soulart Inc.(Soulart公司)

AI总结 Point2Pose通过生成模型和大规模数据集提升3D人体姿态估计的准确性与鲁棒性。

Comments WACV 2026 camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09383 2025-12-12 cs.CV

Perception-Inspired Color Space Design for Photo White Balance Editing

受感知启发的色彩空间设计用于照片白平衡编辑

Yang Cheng, Ziteng Cui, Shenghan Su, Lin Gu, Zenghui Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) The University of Tokyo(东京大学) Tohoku University(东北大学)

AI总结 本文提出了一种基于受感知启发的可学习HSI色彩空间的白平衡校正框架,以解决传统加色模型在复杂光照条件下的局限性。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07730 2025-12-12 cs.CV cs.AI

SAVE: Sparse Autoencoder-Driven Visual Information Enhancement for Mitigating Object Hallucination

SAVE:基于稀疏自编码器的视觉信息增强用于缓解物体幻觉

Sangha Park, Seungryong Yoo, Jisoo Mok, Sungroh Yoon

机构 * Department of Electrical and Computer Engineering, Seoul National University(电子与计算机工程系,首尔国立大学) Daegu Gyeongbuk Institute of Science and Technology(大邱庆州科学技术院) IPAI, AIIS, ASRI, INMC, and ISRC, Seoul National University(IPAI、AIIS、ASRI、INMC 和 ISRC,首尔国立大学)

AI总结 SAVE通过引导模型沿稀疏自编码器潜在特征减少物体幻觉,提升视觉理解能力。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07702 2025-12-12 cs.CV cs.AI

Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment

引导不应生成的内容:用于文本-图像对齐的自动化负提示

Sangha Park, Eunji Kim, Yeongtak Oh, Jooyoung Choi, Sungroh Yoon

机构 * Department of Electrical and Computer Engineering, Seoul National University(电子与计算机工程系,首尔国立大学) Amazon(亚马逊) IPAI, AIIS, ASRI, INMC, and ISRC, Seoul National University(IPAI、AIIS、ASRI、INMC 和 ISRC,首尔国立大学)

AI总结 本文提出NPC方法,通过自动化负提示生成提升文本-图像对齐效果,在GenEval++和Imagine-Bench上取得优异成绩。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00846 2025-12-12 cs.CV

AFRAgent : An Adaptive Feature Renormalization Based High Resolution Aware GUI agent

AFRAgent:一种基于自适应特征归一化的高分辨率感知GUI代理

Neeraj Anand, Rishabh Jain, Sohan Patnaik, Balaji Krishnamurthy, Mausoom Sarkar

机构 * Media and Data Science Research, Adobe(Adobe媒体与数据科学研究所)

AI总结 AFRAgent通过自适应特征归一化技术,在保持高分辨率细节的同时,实现了更高效的GUI自动化性能,比现有模型小四分之一。

Comments Accepted at WACV 2026 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00087 2025-12-12 cs.CV

Exploring Automated Recognition of Instructional Activity and Discourse from Multimodal Classroom Data

探索多模态课堂数据中教学活动和话语的自动化识别

Ivo Bueno, Ruikun Hou, Babette Bühler, Tim Fütterer, James Drimalla, Jonathan Kyle Foster, Peter Youngs, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci

AI总结 本文通过多模态分析方法,实现了课堂活动中教学活动和话语的自动化识别,展示了微调模型在视频和 transcripts 上的高准确率,为可扩展的教师反馈系统提供了基础。

Comments This article has been accepted for publication in the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21888 2025-12-12 cs.CV

CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding

CAPE:一种基于CLIP的互补热图线索点集用于具身参照理解

Fevziye Irem Eyiokur, Dogucan Yaman, Hazım Kemal Ekenel, Alexander Waibel

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Istanbul Technical University(伊斯坦布尔技术大学) Carnegie Mellon University(卡内基梅隆大学) KIT Campus Transfer GmbH (KCT)(KIT校园转移有限责任公司)

AI总结 CAPE通过双模型框架和CLIP-aware Pointing Ensemble模块,提升具身参照理解任务中指向线索的多模态推理能力,实现75.0 mAP的高精度表现。

Comments Accepted by WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13901 2025-12-12 cs.CV

Dressing the Imagination: A Dataset for AI-Powered Translation of Text into Fashion Outfits and A Novel NeRA Adapter for Enhanced Feature Adaptation

为想象力着装:一个用于人工智能驱动的文本到时尚装扮的数据库及一种新的NeRA适配器用于增强特征适应

Gayatri Deshmukh, Somsubhra De, Chirag Sehgal, Jishu Sen Gupta, Sparsh Mittal

机构 * IIT Madras(印度理工学院Madras分校) Delhi Technological University(德里技术大学) IIT BHU(印度理工学院BHU分校) IIT Roorkee(印度理工学院Roorkee分校)

AI总结 FLORA数据集和NeRA适配器旨在提升人工智能生成时尚设计的精度与风格丰富度。

Comments Accepted as a Conference Paper at WACV 2026 (USA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10102 2025-12-12 cs.CV

Hierarchical Instance Tracking to Balance Privacy Preservation with Accessible Information

层级实例跟踪以平衡隐私保护与可获取信息

Neelima Prasad, Jarek Reynolds, Neel Karsanbhai, Tanusree Sharma, Lotus Zhang, Abigale Stangl, Yang Wang, Leah Findlater, Danna Gurari

机构 * University of Colorado Boulder(科罗拉多大学博尔德分校) Pennsylvania State University(宾夕法尼亚州立大学) University of Washington(华盛顿大学) Georgia Institute of Technology(佐治亚理工学院) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 本文提出层级实例跟踪任务,构建首个支持该任务的基准数据集,通过评估多种模型展示数据集的挑战性,旨在平衡隐私保护与信息可获取性。

Comments Accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏