arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高校专区

Huazhong University of Science and Technology(华中科技大学)

2026-02-12 至 2026-02-12 共收录 5
2602.11007 2026-02-12 cs.CV

LaSSM: Efficient Semantic-Spatial Query Decoding via Local Aggregation and State Space Models for 3D Instance Segmentation

LaSSM:通过局部聚合和状态空间模型实现高效的语义-空间查询解码用于3D实例分割

Lei Yao, Yi Wang, Yawen Cui, Moyun Liu, Lap-Pui Chau

机构 * Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University(电子工程系,香港理工大学) School of Mechanical Science and Engineering, Huazhong University of Science and Technology(机械科学与工程学院,华中科技大学)

AI总结 LaSSM通过局部聚合和状态空间模型实现高效语义-空间查询解码,提升3D实例分割性能。

Comments Accepted at IEEE-TCSVT

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02172 2026-02-12 cs.CV cs.AI cs.MM

GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting

GaussianCross: 通过高斯点撒技术实现跨模态自监督3D表示学习

Lei Yao, Yi Wang, Yi Zhang, Moyun Liu, Lap-Pui Chau

机构 * Hong Kong Polytechnic University(香港理工大学) Huazhong University of Science and Technology(华中科技大学)

AI总结 GaussianCross通过高斯点撒技术实现跨模态自监督3D表示学习,提升3D点云的表示质量和泛化能力。

Comments 14 pages, 8 figures, accepted by MM'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10575 2026-02-12 cs.CV cs.AI cs.CY

MetaphorStar: Image Metaphor Understanding and Reasoning with End-to-End Visual Reinforcement Learning

MetaphorStar: 基于端到端视觉强化学习的图像隐喻理解与推理

Chenhao Zhang, Yazhe Niu, Hongsheng Li

机构 * Shanghai AI Laboratory(上海人工智能实验室) Huazhong University of Science and Technology(华中科技大学) The Chinese University of Hong Kong MMLab(香港中文大学MMLab)

AI总结 MetaphorStar通过端到端视觉强化学习框架提升图像隐喻理解与推理能力,显著优于现有模型。

Comments 14 pages, 4 figures, 11 tables; Code: https://github.com/MING-ZCH/MetaphorStar, Model & Dataset: https://huggingface.co/collections/MING-ZCH/metaphorstar

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00891 2026-02-12 cs.CV

Accelerating Streaming Video Large Language Models via Hierarchical Token Compression

通过分层令牌压缩加速流式视频大型语言模型

Yiyu Wang, Xuyang Liu, Xiyan Gui, Xinying Lin, Boxue Yang, Chenfei Liao, Tailai Chen, Linfeng Zhang

机构 * EPIC Lab, Shanghai Jiao Tong University(上海交通大学EPIC实验室) Sichuan University(四川大学) Huazhong University of Science and Technology(华中科技大学) Sun Yat-sen University(中山大学) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

AI总结 STC通过分层令牌压缩技术,显著降低流式视频大语言模型的处理延迟,同时保持高精度。

Comments Code is avaliable at \url{https://github.com/lern-to-write/STC}

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09221 2026-02-12 cs.HC cs.LG

CMCRD: Cross-Modal Contrastive Representation Distillation for Emotion Recognition

CMCRD:跨模态对比表示蒸馏用于情绪识别

Siyuan Kan, Huanyu Wu, Zhenyao Cui, Fan Huang, Xiaolong Xu, Dongrui Wu

机构 * Key Laboratory of Image Processing and Intelligent Control, Ministry of Education, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(图像处理与智能控制重点实验室、教育部、自动化学院、华中科技大学) Wuhan Institute of Digital Engineering(武汉数字工程研究所) Shanghai Jiao Tong University(上海交通大学)

AI总结 CMCRD通过跨模态对比表示蒸馏提升情绪识别准确性,减少多模态数据需求,实验表明在EEG和眼动数据间相互辅助训练可提高分类准确率约6.2%。

Journal ref IEEE Trans. on Affective Computing, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏