arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision

2026-06-10 至 2026-06-10 共收录 6
2606.05399 2026-06-10 cs.CV 版本更新

UniPixie: Unified and Probabilistic 3D Physics Learning via Flow Matching

UniPixie: 基于流匹配的统一概率三维物理学习

Qilin Huang, Quynh Anh Huynh, Long Le, Chen Wang, Chuhao Chen, Ryan Lucas, Eric Eaton, Lingjie Liu

机构 * University of Pennsylvania(宾夕法尼亚大学) Southern University of Science and Technology(南方科技大学) MIT(麻省理工学院)

AI总结 提出UniPixie框架,通过流匹配学习从单张视觉输入到连续可控材料属性分布的映射,实现多样物理场生成并降低杨氏模量预测误差超50%。

Comments Published at CVPR 2026 as a Highlight. Project page: https://unipixie.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13776 2026-06-10 cs.CY cs.CL cs.CR cs.CV 版本更新

Who Gets Flagged? The Pluralistic Evaluation Gap in AI Content Watermarking

谁被标记?AI内容水印中的多元评估差距

Alexander Nemecek, Osama Zafar, Yuqiao Xu, Wenbiao Li, Erman Ayday

机构 * Case Western Reserve University(凯斯西储大学)

AI总结 本文揭示AI内容水印在不同语言、文化和群体间存在系统性偏差,提出跨语言检测一致性、文化多样性覆盖和检测指标人口统计分解三个评估维度,主张水印部署前必须进行公平性审计。

Comments 7 pages. Accepted at the Multimodal Alignment for a Pluralistic Society (MAPS) Workshop, CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12675 2026-06-10 cs.CV cs.AI 版本更新

Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling

Scone:通过统一理解-生成建模弥合主体驱动图像生成中的组合与区分

Yuran Wang, Bohan Zeng, Chengzhuo Tong, Wenxuan Liu, Yang Shi, Xiaochen Ma, Hao Liang, Yuanxing Zhang, Wentao Zhang

机构 * Peking University(北京大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队) Zhongguancun Academy(中关村学院) HKUST(香港科技大学) Beijing Key Laboratory of Data Intelligence and Security (Peking University)(北京数据智能与安全重点实验室(北京大学))

AI总结 提出Scone方法,通过统一理解-生成模型结合组合与区分能力,采用两阶段训练实现主体身份保持与干扰最小化,在双基准上优于现有开源模型。

Comments CVPR 2026 Highlight. Code: https://github.com/Ryann-Ran/Scone

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23499 2026-06-10 cs.RO cs.AI 版本更新

TaCarla: A comprehensive benchmarking dataset for end-to-end autonomous driving

TaCarla: 端到端自动驾驶的综合基准数据集

Tugrul Gorgulu, Atakan Dag, M. Esat Kalfaoglu, Halil Ibrahim Kuru, Baris Can Cam, Halil Ibrahim Ozturk, Ozsel Kilinc

机构 * Tuğrul Gorgülü *†(土耳其巴伊塞蒂大学) Atakan Dağ †(土耳其巴伊塞蒂大学) M. Esat Kalfaoğlu ‡(土耳其巴伊塞蒂大学) Halil İbrahim Kuru †(土耳其巴伊塞蒂大学) Barış Can Cam †(土耳其巴伊塞蒂大学) Halil İbrahim Öztürk †(土耳其巴伊塞蒂大学) Özsel Kılınç §(土耳其巴伊塞蒂大学)

AI总结 针对现有自动驾驶数据集不完整、行为多样性不足及闭环评估缺失等问题,基于CARLA Leaderboard 2.0挑战场景收集超过285万帧的多任务数据集,支持规划、检测、预测及视觉语言动作模型,并提供数值稀有度评分。

Comments Accepted at the Third Workshop on Simulation for Autonomous Driving (SAD), CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20850 2026-06-10 cs.CV cs.RO 版本更新

Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves

Glove2Hand:从多模态传感手套合成自然的手-物体交互

Xinyu Zhang, Ziyi Kou, Chuan Qin, Mia Huang, Ergys Ristani, Ankit Kumar, Lele Chen, Kun He, Abdeslam Boularias, Li Guan

机构 * Meta Reality Labs(Meta现实实验室) Rutgers University(罗格斯大学)

AI总结 提出Glove2Hand框架,将多模态传感手套视频转化为逼真的裸手,并保留物理交互动态;引入3D高斯手模型和扩散手恢复器,创建HandSense数据集,提升下游任务性能。

Comments CVPR 2026 Highlight. This version includes the motion retarget process in the appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03118 2026-06-10 cs.CV cs.AI 版本更新

NuWa: Deriving Lightweight Class-Specific Vision Transformers for Edge Devices

NuWa: 为边缘设备导出轻量级类别特定视觉Transformer

Ziteng Wei, Qiang He, Bing Li, Feifei Chen, Hai Jin, Yun Yang

机构 * National Engineering Research Center for Big Data Technology and System, Services Computing Technology and System Lab, Cluster and Grid Computing Lab(大数据技术与系统国家工程研究中心、服务计算技术与系统实验室、集群与网格计算实验室) Swinburne University of Technology(斯威本科技大学) Deakin University(迪金大学)

AI总结 针对边缘设备只需识别特定类别的问题,提出NuWa方法,通过自知识净化去除有害权重,并利用闭式优化高效导出紧凑ViT,无需重训练即可提升类别精度并加速推理。

Comments Accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏