arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

ACM International Conference on Multimedia · 会议 · Multimedia

共收录 1940
2605.26941 2026-05-27 cs.IR cs.MM

The 2nd EReL@MIR Workshop on Efficient Representation Learning for Multimodal Information Retrieval

第二届EReL@MIR研讨会:面向多模态信息检索的高效表示学习

Junchen Fu, Xuri Ge, Xin Xin, Alexandros Karatzoglou, Ioannis Arapakis, Xi Wang, Qijiong Liu, Qian Li, Joemon M. Jose

AI总结 本研讨会旨在探讨多模态基础模型在信息检索中的效率瓶颈,并提出通过组织学术与工业界交流,推动高效表示学习的新方法、度量标准和基准。

Comments Accepted as a workshop proposal at ACM Multimedia 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26313 2026-05-27 cs.MM

Reproducibility Companion Paper: Swarical: An Integrated Hierarchical Approach to Localizing Flying Light Specks

可重复性配套论文:Swarical:一种用于定位飞行光点的集成层次化方法

Hamed Alimohammadzadeh, Shahram Ghandeharizadeh, Federico Cunico, Joshua Springer

AI总结 本论文提供可重复性实验的工件和指南,验证ACM Multimedia 2024论文中提出的Swarical层次化定位技术,该技术使微型无人机(飞行光点)能准确高效地定位和照亮复杂2D和3D形状。

Comments Reproducibility is one of the foundations of reliable science and engineering. This paper establishes the reproducibility of the Swarical decentralized technique by colleagues in Italy and Iceland. Appeared in Proceedings of the 33rd ACM International Conference on Multimedia (MM '25), October 27-31, 2025, Dublin, Ireland. ACM, New York, NY, USA, 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23774 2026-05-25 cs.MM

Swarical: An Integrated Hierarchical Approach to Localizing Flying Light Specks

Swarical:一种集成层次化方法用于定位飞行光点

Hamed Alimohammadzadeh, Shahram Ghandeharizadeh

AI总结 提出一种基于群体智能的层次化定位技术Swarical,通过异构传感器配置和点云转换,使微型无人机FLS能高精度高效定位并照亮复杂2D/3D形状,相比现有技术速度提升2倍以上。

Comments Appeared in proceedings of the 32nd ACM International Conference on Multimedia (MM '24), October 28-November 1, 2024, Melbourne, VIC, Australia. ACM, New York, NY, USA, 9 pages. Source code available at: https://github.com/flyinglightspeck/Swarical. See https://youtu.be/NHMGT-Pjy-A for a demonstration

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05269 2026-05-13 cs.CV

B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding

B4DL:用于空间时间理解的4D激光雷达LLM基准

Changho Choi, Youngwoo Shin, Gyojin Han, Dong-Jae Lee, Junmo Kim

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院)

AI总结 本文提出B4DL基准,用于训练和评估多模态大语言模型对4D激光雷达数据的理解,结合可扩展的数据生成管道和新模型,实现动态户外环境的空间时间推理。

Comments Accepted at ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07489 2026-05-11 cs.SD cs.MM eess.SP

A Decomposed Retrieval-Edit-Rerank Framework for Chord Generation

为和弦生成设计的分解检索-编辑-重排序框架

Qiqi He, Dichucheng Li, Xiaoheng Sun, Anqi Huang

机构 * Individual Researcher(个人研究者)

AI总结 本文提出分解检索-编辑-重排序框架,通过分阶段处理提升和弦生成的多样性与音乐理论可行性平衡能力。

Comments Accepted by the 2026 ACM International Conference on Multimedia Retrieval (ICMR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27620 2026-05-01 cs.CV

SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation

SpaAct:基于课程适应的空间激活转换学习用于视觉语言导航

Pengna Li, Kangyi Wu, Shaoqing Xu, Fang Li, Hanbing Li, Lin Zhao, Kailin Lyu, Long Chen, Zhi-Xin Yang, Nanning Zheng

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(人机混合增强智能国家重点实验室) National Engineering Research Center for Visual Information and Applications(视觉信息与应用国家工程研究中心) Institute of Artificial Intelligence and Robotics(人工智能与机器人研究所) The State Key Laboratory of Internet of Things for Smart City(智能城市物联网国家重点实验室) Centre for Artificial Intelligence and Robotics(人工智能与机器人中心) University of Macau(澳门大学) Xiaomi EV(小i EV) School of Automation(自动化学院) Beijing Institute of Technology(北京理工大学) Institute of Automation(自动化研究所) Chinese Academy of Sciences(中国科学院)

AI总结 SpaAct通过引入空间激活任务和课程适应方法,提升视觉语言导航中动态空间感知能力,实现更高效的导航性能。

Comments Submmited to ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23584 2026-04-28 cs.CV cs.IR

Identity-Decoupled Anonymization for Visual Evidence in Multi-modal Retrieval-Augmented Generation

多模态检索增强生成中视觉证据的身份解耦匿名化

Zehua Cheng, Wei Dai, Jiahao Sun

机构 * Department of Computer Science University of Oxford, UK

AI总结 本文提出身份解耦MRAG框架,通过在检索与生成之间插入生成匿名化模块,解耦人脸身份与属性,利用生成对抗网络和拒绝采样器实现隐私保护。

Comments ACM International Conference on Multimedia Retrieval 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23551 2026-04-28 cs.CV

Spatiotemporal Degradation-Aware 3D Gaussian Splatting for Realistic Underwater Scene Reconstruction

时空降质感知的3D高斯点散布用于逼真水下场景重建

Shaohua Liu, Ning Gao, Zuoya Gu, Hongkun Dou, Yue Deng, Hongjue Li

机构 * School of Astronautics Beihang University Beijing China(航天学院 北航 北京 中国) School of Artificial Intelligence Beihang University Beijing China(人工智能学院 北航 北京 中国) Zhongguancun Academy Beijing China(中关村学院 北京 中国) Beihang University School of Astronautics Beijing China(北航 航天学院 北京 中国) Beihang University(北航) Zhongguancun Academy(中关村学院)

AI总结 本文提出MarineSTD-GS框架,通过建模时空降质实现逼真水下场景重建,引入内在高斯和降质高斯,结合时空降质建模模块,改进几何和外观估计,实验表明其在处理时空降质和合成新视角方面优于现有方法。

Comments 12 pages, 10 figures, 6 tables. Author version of the paper published in Proceedings of ACM Multimedia 2025

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia (ACM MM 2025), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16207 2026-04-28 cs.CV cs.AI

AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection

AIFIND:面向增量面部伪造检测的 artifact 意识细粒度对齐

Hao Wang, Beichen Zhang, Yanpei Gong, Shaoyi Fang, Zhaobo Qi, Yuanrong Xu, Xinyan Liu, Weigang Zhang

机构 * School of Computer Science and Technology, Harbin Institute of Technology, Weihai(哈尔滨工业大学计算机科学与技术学院)

AI总结 本文提出AIFIND,通过语义锚点稳定增量学习,设计Artifact-Driven Semantic Prior Generator和Artifact-Probe Attention模块,解决增量面部伪造检测中的特征漂移和灾难性遗忘问题。

Comments Accepted by ACM International Conference on Multimedia Retrieval (ICMR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20746 2026-04-23 cs.MM

Realistic Virtual Flood Experience System Using 360° Videos and 3D City Models Constructed from Building Footprints

基于360度视频和建筑轮廓构建的3D城市模型的现实虚拟洪水体验系统

Tatsuro Banno, Koki Kawada, Mizuki Takenawa, Masatoshi Denda, Kiyoharu Aizawa

AI总结 本文提出一种整合360度视频与自动构建的3D城市模型的虚拟洪水体验框架,通过建筑轮廓扩展和空间对齐实现逼真洪水可视化,用于提升洪水应急疏散场景的认知。

Comments Accepted by ACM International Conference on Multimedia Retrieval (ICMR 2026), Demonstration

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19689 2026-04-22 cs.AI

A-MAR: Agent-based Multimodal Art Retrieval for Fine-Grained Artwork Understanding

A-MAR:基于代理的多模态艺术检索用于细粒度艺术作品理解

Shuai Wang, Hongyi Zhu, Jia-Hong Huang, Yixian Shen, Chengxi Zeng, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg, Marcel Worring

机构 * University of Amsterdam(阿姆斯特丹大学) University of Bristol(布里斯托大学) Amazon AGI(亚马逊人工智慧) College of Business and Economics(商学院和经济学学院) University of Johannesburg(约翰内斯堡大学)

AI总结 本文提出A-MAR框架,通过结构化推理计划显式指导多模态艺术检索,提升解释质量和证据 grounding 能力,在艺术领域验证了代理式多模态推理的有效性。

Journal ref ICMR 2026, ACM International Conference on Multimedia Retrieval

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05116 2026-04-21 cs.LG cs.AI

FedBCD:Communication-Efficient Accelerated Block Coordinate Gradient Descent for Federated Learning

FedBCD:用于联邦学习的通信高效加速块坐标梯度下降法

Junkang Liu, Fanhua Shang, Yuanyuan Liu, Hongying Liu, Yuangang Li, YunXiang Gong

机构 * School of Artificial Intelligence Xi'an Xidian University, China College of Intelligence Computing Tianjin Tianjin University, China Medical College, Tianjin University\ Cheng Laboratory Tianjin China University of Southern California Los Angeles US School of Artificial Intelligence Medical College, Tianjin University\ Cheng Laboratory University of Southern California

AI总结 本文提出FedBCGD方法,通过划分模型参数为多个块以降低通信开销,并开发FedBCGD+算法实现加速,首次在大规模深度模型中应用参数块通信。

Journal ref ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23827 2026-04-21 cs.LG cs.AI

FedNSAM:Consistency of Local and Global Flatness for Federated Learning

FedNSAM:联邦学习中局部与全局平坦度的一致性

Junkang Liu, Fanhua Shang, Yuxuan Tian, Hongying Liu, Yuanyuan Liu

机构 * College of Intelligence Computing, Tianjin University Tianjin China College of Management Economics, Tianjin University Tianjin China Medical College, Tianjin University\ Cheng Laboratory Tianjin China School of Artificial Intelligence, Xidian University Xian China Computing, Tianjin University Economics, Tianjin University Medical College, Tianjin University\ Cheng Laboratory School of Artificial Intelligence, Xidian University

AI总结 FedNSAM通过引入全局Nesterov动量提升SAM效果,解决高数据异质性下局部平坦度不保证全局模型平坦度的问题,理论证明收敛界更紧,实验验证性能优越。

Journal ref ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19130 2026-04-17 cs.MM

Dual-Stream Decoupled Learning for Temporal Consistency and Speaker Interaction in AVSD

双流解耦学习用于AVSD中的时间一致性和说话人交互

Junhao Xiao, Shun Feng, Zhiyu Wu, Jinghan Yu, Haibiao Yao, Zhiyuan Ma, Jianjun Li, Youjun Bao, Yi Chen

AI总结 本文提出D$^2$Stream双流框架,通过解耦时间连续性和人际交互任务,提升AVSD性能,实现95.6%的mAP和更广的泛化能力。

Comments Submitted to ACMMM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14779 2026-04-17 cs.CV cs.CL

AIM: Asymmetric Information Masking for Visual Question Answering Continual Learning

AIM:面向视觉问答持续学习的非对称信息掩码

Peifeng Zhang, Zice Qiu, Donghua Yu, Shilei Cao, Juepeng Zheng, Yutong Lu, Haohuan Fu

机构 * Sun Yat-Sen University(中山大学) National Supercomputing Center in Shenzhen(深圳国家超算中心) Tsinghua University(清华大学)

AI总结 针对视觉问答持续学习中因模型结构不对称导致的灾难性遗忘问题,提出AIM方法,通过模态特定敏感性指导的针对性掩码平衡稳定性与可塑性,提升在VQA v2和GQA任务中的平均性能和平均遗忘度。

Comments 18 pages, 9 figures. Submitted to ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12650 2026-04-15 cs.CV cs.MM

Listening Deepfake Detection: A New Perspective Beyond Speaking-Centric Forgery Analysis

监听深度伪造检测:一种超越以说话为中心的伪造分析的新视角

Miao Liu, Fangda Wei, Jing Wang, Xinyuan Qian

机构 * Beijing Institute of Technology(北京理工大学) University of Science and Technology Beijing(北京科技大学) University of Science(科学技术大学)

AI总结 本文提出监听深度伪造检测任务,引入首个专门为此任务设计的ListenForge数据集,并提出MANet网络以捕捉监听伪造中的细微运动不一致,实验表明现有说话深度伪造检测模型在监听场景中表现不佳。

Comments Submitted to ACMMM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12315 2026-04-15 cs.CV cs.MM

GTPBD-MM: A Global Terraced Parcel and Boundary Dataset with Multi-Modality

GTPBD-MM:一种具有多模态的全球梯田地块和边界数据集

Zhiwei Zhang, Xingyuan Zeng, Xinkai Kong, Kunquan Zhang, Haoyuan Liang, Bohan Shi, Juepeng Zheng, Jianxi Huang, Yutong Lu, Haohuan Fu

机构 * Sun Yat-sen University(中山大学) Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院) China Agricultural University(中国农业大学) Southwest Jiaotong University(西南交通大学) Northeastern University(东北大学) National Supercomputing Center in Shenzhen(深圳国家超算中心)

AI总结 本文提出GTPBD-MM数据集,用于复杂梯田地块提取,结合高分辨率影像、文本描述和DEM数据,提出ETTerra模型,通过文本和地形信息提升提取精度。

Comments 15 pages, 11 figures. Submitted to ACM Multimedia 2026 Dataset Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24168 2026-04-15 cs.AI

MGA: Memory-Driven GUI Agent for Observation-Centric Interaction

MGA:基于记忆的GUI代理用于以观察为中心的交互

Weihua Cheng, Junming Liu, Yifei Sun, Botian Shi, Yirong Chen, Ding Wang

机构 * Shanghai Tech University(上海科技大学) Tongji University(同济大学) East China University of Science and Technology(东华大学) Shanghai AI Laboratory(上海人工智能实验室)

AI总结 MGA通过解耦长周期轨迹和结构化状态记忆机制,减少认知负担和系统复杂度,实现高效的GUI自动化。

Comments Submitted to ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04672 2026-04-14 cs.CV cs.CL

Agri-R1: Agricultural Reasoning for Disease Diagnosis via Automated-Synthesis and Reinforcement Learning

Agri-R1:通过自动合成与强化学习进行农业疾病诊断的推理增强

Wentao Zhang, Mingkun Xu, Qi Zhang, Shangyang Li, Derek F. Wong, Lifei Wang, Yanchao Yang, Lina Lu, Tao Fang

机构 * Shandong University of Technology(山东理工大学) Guangdong Institute of Intelligence Science and Technology(广东智能科学与技术研究院) Faculty of Data Science, City University of Macau(澳门城市大学数据科学学院) School of Physical Science and Technology, Beijing University of Posts and Telecommunications(北京邮电大学物理科学与技术学院) NLP2CT Lab, Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系自然语言处理与中葡机器翻译实验室) Institute of International Language Services Studies, Macau Millennium College(澳门千禧学院国际语言服务研究所)

AI总结 Agri-R1通过自动合成与强化学习提升农业疾病诊断,利用仅19%的数据生成高质量推理数据,采用改进的奖励函数提升模型在疾病识别和农业知识问答上的性能。

Comments This paper is submitted for review to the 2026 ACM MM Conference. The corresponding authors are Tao Fang and Lina Lu, where Tao Fang is the senior Corresponding Author (Last Author) and the principal supervisor of this work, having led the research design, guided the methodology, and overseen the entire project

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01944 2026-04-13 cs.CV

MobileMold: A Smartphone-Based Microscopy Dataset for Food Mold Detection

MobileMold: 一种基于智能手机的显微图像数据集用于食品霉变检测

Dinh Nam Pham, Leonard Prokisch, Bennet Meyer, Jonas Thumbs

机构 * Technical University of Berlin(柏林工业大学) University of Regensburg(雷根斯堡大学) ETH Zurich(苏黎世联邦理工学院) University of Tübingen(蒂宾根大学)

AI总结 本文提出MobileMold数据集,包含4941张显微图像,用于食品霉变检测与分类。通过多种深度学习模型和增强策略,达到高准确率,验证了数据集在食品变质检测中的有效性。

Comments Accepted to ACM Multimedia Systems (MMSys'26). Dataset and code available at https://mobilemold.github.io/dataset/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08579 2026-04-13 cs.LG cs.AI

On the Spectral Geometry of Cross-Modal Representations: A Functional Map Diagnostic for Multimodal Alignment

关于跨模态表示的谱几何:一种用于多模态对齐的功能图诊断

Krisanu Sarkar

机构 * Indian Institute of Technology Bombay(印度理工学院孟买分校)

AI总结 本文研究了基于功能图框架的跨模态对齐,发现尽管功能图在不同监督预算下表现不如Procrustes对齐,但揭示了多模态表示的结构特性,提出了三种诊断量以表征跨模态表示兼容性。

Comments Under review at ACMMM Brave New Ideas Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08184 2026-04-10 cs.SD cs.AI

AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan

AT-ADD:所有类型音频深度伪造检测挑战评估计划

Yuankun Xie, Haonan Cheng, Jiayi Zhou, Xiaoxuan Guo, Tao Wang, Jian Liu, Weiqiang Wang, Ruibo Fu, Xiaopeng Wang, Hengyan Huang, Xiaoying Huang, Long Ye, Guangtao Zhai

机构 * Communication University of China(中国传媒大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Beijing Institute of Technology(北京理工大学) Shanghai Jiao Tong University(上海交通大学)

AI总结 针对现有音频深度伪造检测方法在鲁棒性和泛化能力上的不足,AT-ADD挑战旨在通过标准化数据集和评估协议提升音频伪造检测的实用性和鲁棒性。

Comments Accepted to the ACM Multimedia 2026 Grand Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05375 2026-04-08 cs.MM

DAT: Dual-Aware Adaptive Transmission for Efficient Multimodal LLM Inference in Edge-Cloud Systems

DAT:双重视觉与带宽感知的自适应传输,用于边缘-云计算系统中高效多模态大语言模型推理

Qi Guo, Zheming Yang, Yunqing Hu, Chang Zhao, Wen Ji

AI总结 本文提出DAT方法,通过轻量级边缘模型过滤非目标帧并触发MLLM推理,结合高效微调策略和多流自适应传输优化,实现高效多模态大语言模型推理,提升语义生成质量与低延迟警报性能。

Comments 10 pages, 6 figures. Submitted to ACM Multimedia 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27781 2026-03-31 cs.CV

GS3LAM: Gaussian Semantic Splatting SLAM

GS3LAM: 基于高斯语义散射的SLAM

Linfei Li, Lin Zhang, Zhong Wang, Ying Shen

机构 * School of Software Engineering, Tongji University(同济大学软件工程学院) Department of Automation, Shanghai Jiaotong University(上海交通大学自动化系)

AI总结 GS3LAM通过多模态数据融合实现实时一致的语义密集地图,采用语义高斯场和多模态误差约束联合优化,引入深度自适应尺度正则化和随机采样关键帧映射策略,提升跟踪鲁棒性和渲染质量。

Comments Accepted by ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24569 2026-03-31 cs.CV

POLY-SIM: Polyglot Speaker Identification with Missing Modality Grand Challenge 2026 Evaluation Plan

POLY-SIM:基于缺失模态的多模态说话人识别挑战2026评估计划

Marta Moscati, Muhammad Saad Saeed, Marina Zanoni, Mubashir Noman, Rohan Kumar Das, Monorama Swain, Yufang Hou, Elisabeth Andre, Khalid Mahmood Malik, Markus Schedl, Shah Nawaz

机构 * Institute of Computational Perception, Johannes Kepler University Linz, Austria(林茨约翰·开普勒大学计算感知研究所) University of Michigan-Flint, USA(密歇根大学弗林特分校) Sapienza University of Rome, Italy(罗马大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Fortemedia Singapore, Singapore(新加坡Fortemedia公司) IT:U Interdisciplinary Transformation University Austria(奥地利跨学科转型大学) University of Augsburg, Germany(奥格斯堡大学) Human-centered AI Group, AI Lab, Linz Institute of Technology, Austria(奥地利林茨理工学院人工智能实验室人本人工智能组)

AI总结 本文针对多模态说话人识别中缺失模态和跨语言条件下的鲁棒性问题,提出POLY-SIM 2026挑战,设计标准化基准和评估框架以推动更实用的系统发展。

Comments Grand challenge at ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16541 2026-03-30 cs.CV cs.MM

QPT V2: Masked Image Modeling Advances Visual Scoring

QPT V2:掩码图像建模在视觉评分中的进展

Qizhi Xie, Kun Yuan, Yunpeng Qu, Mingda Wu, Ming Sun, Chao Zhou, Jihong Zhu

机构 * Tsinghua University(清华大学) Kuaishou Technology(快手科技)

AI总结 本文提出QPT V2框架,通过掩码图像建模提升视觉质量和审美评分能力,通过数据筛选、降质引入和模型结构调整,在11个下游基准测试中表现优异。

Comments 8 pages, 6 figures. Accepted by ACM MM 24

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07315 2026-03-25 cs.MA cs.AI cs.AR

VLM-CAD: VLM-Optimized Collaborative Agent Design Workflow for Analog Circuit Sizing

VLM-CAD:面向模拟电路尺寸设计的VLM优化协作代理设计流程

Guanyuan Pan, Shuai Wang, Yugui Lin, Tiansheng Zhou, Pietro Liò, Zhenxin Zhao, Yaqi Wang

机构 * Hangzhou Dianzi University(杭州电子科技大学) SCBC, Guangdong University of Foreign Studies(广东外语外贸大学) University of Cambridge(剑桥大学)

AI总结 本文提出VLM-CAD流程,通过神经符号结构解析模块和ExTuRBO方法提升多模态推理的鲁棒性和可解释性,实验表明其在模拟电路设计中具有高准确性和低功耗。

Comments submitted to the 34th ACM International Conference on Multimedia (ACMMM 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16344 2026-03-25 cs.CV cs.AI

SOAP: Enhancing Spatio-Temporal Relation and Motion Information Capturing for Few-Shot Action Recognition

SOAP:增强少样本动作识别中的时空关系和运动信息捕捉

Wenbo Huang, Jinghui Zhang, Xuwei Qian, Zhen Wu, Meng Wang, Lei Zhang

机构 * Southeast University(东南大学) Tongji University(同济大学) Nanjing Normal University(南京师范大学)

AI总结 本文提出SOAP架构,通过融合时空特征与多帧帧元组提升少样本动作识别性能,实现更全面的运动信息捕捉,优于传统方法。

Comments Accepted by ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24133 2026-03-17 cs.CV

FocusTrack: One-Stage Focus-and-Suppress Framework for 3D Point Cloud Object Tracking

FocusTrack: 一种用于3D点云目标跟踪的一阶段聚焦与抑制框架

Sifan Zhou, Jiahao Nie, Ziyu Zhao, Yichao Cao, Xiaobo Lu

AI总结 FocusTrack通过IMM和Focus-and-Suppress Attention实现运动-语义联合建模,以一阶段框架在KITTI等基准上达到SOTA性能并实现105 FPS高帧率。

Comments Acceptted in ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00782 2026-03-17 cs.GR cs.AI cs.CV cs.MM cs.SD eess.AS

SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video Generation

SpA2V: 利用空间听觉线索进行音频驱动的空间感知视频生成

Kien T. Pham, Yingqing He, Yazhou Xing, Qifeng Chen, Long Chen

AI总结 SpA2V通过利用空间听觉线索生成与输入音频在语义和空间上一致的视频,改进了现有方法对语义信息的依赖,提出两阶段框架实现视频场景布局生成与生成。

Comments The 33rd ACM Multimedia Conference (MM '25)

详情

展开后加载摘要…

URL PDF HTML 收藏