arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision

2026-07-14 至 2026-07-14 共收录 5
2606.01063 2026-07-14 cs.AI 版本更新

MindClaw: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention

MindClaw: 用于精确干预的闭环具身心理状态推理

Ruoxuan Zhang, Qiaoqiao Wan, Zhengguang Wang, Chenghao Yu, Hongxia Xie, Jianlong Fu, Wen-Huang Cheng

机构 * Jilin University(吉林大学) Microsoft Asia(微软亚洲) National Taiwan University(国立台湾大学)

AI总结 提出MindClaw框架,通过闭环具身心理状态推理实现精确干预,结合多源输入、信念记忆、认知触发技能和动作生成,在动态环境中优化干预时机。

Comments Extended version of the CVPR 2026 paper *MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents*. This work is in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13840 2026-07-14 cs.CV 版本更新

MoLingo: Motion-Language Alignment for Text-to-Human Motion Generation

MoLingo:用于文本到运动生成的运动-语言对齐

Yannan He, Garvita Tiwari, Xiaohan Zhang, Pankaj Bora, Tolga Birdal, Jan Eric Lenssen, Gerard Pons-Moll

机构 * University of Tübingen(图宾根大学) Tübingen AI Center(图宾根人工智能中心) Max Planck Institute for Informatics(马克斯·普朗克信息学研究所) Imperial College London(伦敦帝国理工学院) Zuse School ELIZA(Zuse ELIZA 学院)

AI总结 本文提出MoLingo模型,通过在连续潜在空间中去噪生成逼真的人体运动。研究如何构建语义对齐的潜在空间和最佳注入文本条件以提高运动真实性与描述一致性。

Comments Accepted by CVPR 2026. Project page: https://hynann.github.io/molingo/MoLingo.html. Title type fixed, content unchanged

Journal ref Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recogn. (CVPR), 2026, pp. 38387-38398

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22042 2026-07-14 cs.CV cs.AI 版本更新

Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models

基于部分-整体语义代表性与不确定性引导的组合对齐超几何视觉语言模型

Hayeon Kim, Ji Ha Jang, Junghun James Kim, Se Young Chun

机构 * Dept. of Electrical and Computer Engineering(电子与计算机工程系) INMC & IPAI(INMC与IPAI) Seoul National University(首尔国立大学)

AI总结 本文提出UNCHA方法,通过超几何不确定性建模部分-整体语义代表性,提升超几何VLM对多物体场景的组合结构理解能力,实现零样本分类等任务的最先进性能。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21511 2026-07-14 cs.CV 版本更新

Back to Point: Exploring Point-Language Models for Zero-Shot 3D Anomaly Detection

回到点:探索用于零样本3D异常检测的点语言模型

Kaiqiang Li, Gang Li, Mingle Zhou, Min Li, Delong Han, Jin Wan

机构 * Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences), Jinan, China(计算机功率网络与信息安全重点实验室,教育部,山东计算机科学中心(济南国家超级计算机中心),齐鲁大学(山东科学院),济南,中国) Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science, Jinan, China(山东省计算功率互联网与服务计算重点实验室,山东省计算机科学基础研究中心,济南,中国)

AI总结 本文提出BTP框架,通过结合3D点云和文本嵌入,提升零样本3D异常检测的性能,实验表明其在Real3D-AD和Anomaly-ShapeNet上表现优异。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19311 2026-07-14 cs.CV cs.AI 版本更新

MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture

MixFlow训练:用慢插值混合减轻曝光偏差

Hui Li, Fu-Yun Wang, Haoyuan Xia, Jiayue Lyu, Kaihui Cheng, Siyu Zhu, Jingdong Wang

AI总结 研究扩散模型训练-测试差异问题,提出MixFlow方法,利用慢流现象的慢插值混合对预测网络后处理,在类条件图像生成和文本到图像生成实验中验证有效,在RAE模型的ImageNet生成任务上取得好结果。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏