arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

ACM International Conference on Multimedia · 会议 · Multimedia

共收录 1940
2506.05696 2025-10-31 cs.CV

MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory

Ana Carolina Condez, Diogo Tavares, João Magalhães

机构 * NOVA LINCS, NOVA School of Science and Technology(NOVA LINCS,NOVA科学与技术学院)

Comments Updated version: corresponds to the ACM MM '25 published paper and includes full appendix material

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14334 2025-10-31 cs.CV cs.HC cs.LG

Evaluating the Evaluators: Towards Human-aligned Metrics for Missing Markers Reconstruction

Taras Kucherenko, Derek Peristy, Judith Bütepage

机构 * SEED, Electronic Arts(SEED,电子艺名)

Comments Accepted at the ACM International Conference on Multimedia 2025 (ACM MM'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01701 2025-10-31 cs.CV

Signal-SGN: A Spiking Graph Convolutional Network for Skeletal Action Recognition via Learning Temporal-Frequency Dynamics

Naichuan Zheng, Yuchen Du, Hailun Xia, Zeyu Liang

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25163 2025-10-30 cs.CV

Target-Guided Bayesian Flow Networks for Quantitatively Constrained CAD Generation

Wenhao Zheng, Chenwei Sun, Wenbo Zhang, Jiancheng Lv, Xianggen Liu

机构 * College of Computer Science, Sichuan University(四川大学计算机科学学院) School of Computer Science and Technology, Xidian University(西安电子科技大学计算机科学与技术学院) Engineering Research Center of Machine Learning and Industry Intelligence, Ministry of Education, Chengdu, China(教育部机器学习与工业智能工程研究中心)

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia (2025) 3330-3339

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22481 2025-10-30 eess.IV cs.AI cs.CV cs.MM

Towards Blind Bitstream-corrupted Video Recovery via a Visual Foundation Model-driven Framework

Tianyi Liu, Kejun Wu, Chen Cai, Yi Wang, Kim-Hui Yap, Lap-Pui Chau

机构 * School of EEE, Nanyang Technological University(南洋理工大学电子工程系) School of EIC, Huazhong University of Science and Technology(华中科技大学电子信息学院) Dept. of EEE, The Hong Kong Polytechnic University(香港理工大学电子工程系)

Comments 10 pages, 5 figures, accepted by ACMMM 2025

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19085 2025-10-30 cs.LG

Clustering-Oriented Generative Attribute Graph Imputation

Mulin Chen, Bocheng Wang, Jiaxin Zhong, Zongcheng Miao, Xuelong Li

机构 * School of Artificial Intelligence, OPtics and ElectroNics (iOPEN)(人工智能学院(iOPEN)) Northwestern Polytechnical University(西北工业大学) Institute of Artificial Intelligence (TeleAI)(人工智能研究所(TeleAI))

Comments Accepted by ACM MM'25

Journal ref ACM MM (2025), pages 1092-1101

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20077 2025-10-29 cs.RO cs.CV cs.HC

Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task Planning

Xun Li, Rodrigo Santa Cruz, Mingze Xi, Hu Zhang, Madhawa Perera, Ziwei Wang, Ahalya Ravendran, Brandon J. Matthews, Feng Xu, Matt Adcock, Dadong Wang, Jiajun Liu

机构 * CSIRO(澳大利亚联邦科学与工业研究组织)

Journal ref MM '25: Proceedings of the 33rd ACM International Conference on Multimedia (2025) Pages 12492 - 12500

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17939 2025-10-29 cs.CV cs.AI

GEMeX-RMCoT: An Enhanced Med-VQA Dataset for Region-Aware Multimodal Chain-of-Thought Reasoning

Bo Liu, Xiangyu Zhao, Along He, Yidi Chen, Huazhu Fu, Xiao-Ming Wu

机构 * The Hong Kong Polytechnic University(香港理工大学) Shenzhen University(深圳大学) West China Hospital of Sichuan University(四川大学华西医院) IHPC, Agency for Science, Technology and Research(科技研究局IHPC)

Comments Accepted at ACM MM 2025 (also known as GEMeX-ThinkVG)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09193 2025-10-28 eess.IV cs.CV

BiECVC: Gated Diversification of Bidirectional Contexts for Learned Video Compression

Wei Jiang, Junru Li, Kai Zhang, Li Zhang

机构 * Bytedance(字节跳动)

Comments Accepted to ACMMM 2025

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia, pp.7248-7257, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22683 2025-10-28 cs.CV

Estimation of Fireproof Structure Class and Construction Year for Disaster Risk Assessment

Hibiki Ayabe, Kazushi Okamoto, Koki Karube, Atsushi Shibata, Kei Harada

机构 * The University of Electro-Communications(电通大学)

Journal ref Workshop on Visual and Signal Communication Technologies in Design of Housing, Urban Spaces, Local Communities, and Human Behavior in conjunction with ACM Multimedia Asia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22513 2025-10-28 cs.LG cs.AI

Toward Robust Signed Graph Learning through Joint Input-Target Denoising

Junran Wu, Beng Chin Ooi, Ke Xu

机构 * National University of Singapore(新加坡国立大学) Zhejiang University(浙江大学) Beihang University(北京航空航天大学)

Comments ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17924 2025-10-28 cs.CV

Gaze into the Heart: A Multi-View Video Dataset for rPPG and Health Biomarkers Estimation

Konstantin Egorov, Stepan Botman, Pavel Blinov, Galina Zubkova, Anton Ivaschenko, Alexander Kolsanov, Andrey Savchenko

机构 * Sber AI Lab(Sber AI实验室) Samara State Medical University(萨马拉州医学大学) ISP RAS Research Center for Trusted Artificial Intelligence(俄罗斯科学院信息与通信技术研究所可信人工智能研究中心)

Comments Accepted to ACMMM 2025, Datasets track

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22828 2025-10-28 cs.CV cs.AI

CapRecover: A Cross-Modality Feature Inversion Attack Framework on Vision Language Models

Kedong Xiu, Sai Qian Zhang

机构 * New York University(纽约大学)

Comments 9 pages, accepted by the 2025 ACM Multimedia Conference. Code is available at https://jus1mple.github.io/Image2CaptionAttack

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17394 2025-10-28 cs.CV cs.AI

HiProbe-VAD: Video Anomaly Detection via Hidden States Probing in Tuning-Free Multimodal LLMs

Zhaolin Cai, Fan Li, Ziwei Zheng, Yanjun Qin

机构 * Xinjiang University(新疆大学) Xi'an Jiaotong University(西安交通大学)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04705 2025-10-28 cs.CV

Identity-Preserving Text-to-Video Generation Guided by Simple yet Effective Spatial-Temporal Decoupled Representations

Yuji Wang, Moran Li, Xiaobin Hu, Ran Yi, Jiangning Zhang, Han Feng, Weijian Cao, Yabiao Wang, Chengjie Wang, Lizhuang Ma

机构 * Shanghai Jiao Tong University, Tencent Youtu Lab(上海交通大学,腾讯云图实验室) Tencent Youtu Lab(腾讯云图实验室) Shanghai Jiao Tong University(上海交通大学) Tencent(腾讯)

Comments ACM Multimedia 2025; code URL: https://github.com/rain152/IPVG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19549 2025-10-28 cs.CV

DEEMO: De-identity Multimodal Emotion Recognition and Reasoning

Deng Li, Bohao Xing, Xin Liu, Baiqiang Xia, Bihan Wen, Heikki Kälviäinen

机构 * Lappeenranta-Lahti University of Technology LUT(拉普兰塔-拉赫蒂技术大学) Nanyang Technological University(南洋理工大学) Brno University of Technology(布拉格技术大学)

Comments Accepted by ACMMM 2025

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04273 2025-10-28 cs.IR cs.CV cs.MM cs.SD eess.AS

Audio Does Matter: Importance-Aware Multi-Granularity Fusion for Video Moment Retrieval

Junan Lin, Daizong Liu, Xianke Chen, Xiaoye Qu, Xun Yang, Jixiang Zhu, Sanyuan Zhang, Jianfeng Dong

机构 * Zhejiang University(浙江大学) Peking University(北京大学) Zhejiang Gongshang University(浙江工商大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) University of Science and Technology of China(中国科学技术大学)

Comments Accepted to ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06278 2025-10-28 cs.RO cs.HC

Robust Understanding of Human-Robot Social Interactions through Multimodal Distillation

Tongfei Bian, Mathieu Chollet, Tanaya Guha

机构 * University of Glasgow(格拉斯哥大学) University of Glasgow School of Computer Science(格拉斯哥大学计算机科学学院)

Comments Accepted by ACM Multimedia 2025, camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20393 2025-10-24 cs.CV cs.MM

Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval

Qing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng Lim

机构 * Singapore Management University(新加坡管理大学)

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20327 2025-10-24 cs.LG cs.AI

LEGO: A Lightweight and Efficient Multiple-Attribute Unlearning Framework for Recommender Systems

Fengyuan Yu, Yuyuan Li, Xiaohua Feng, Junjie Fang, Tao Wang, Chaochao Chen

机构 * Zhejiang University(浙江大学) Hangzhou Dianzi University(杭州电子科技大学)

Comments Accepted by ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19451 2025-10-23 cs.CV cs.MM

Reasoning Like Experts: Leveraging Multimodal Large Language Models for Drawing-based Psychoanalysis

Xueqi Ma, Yanbei Jiang, Sarah Erfani, James Bailey, Weifeng Liu, Krista A. Ehinger, Jey Han Lau

机构 * The University of Melbourne(墨尔本大学) China University of Petroleum (East China)(中国石油大学(华东))

Comments Accepted by ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02266 2025-10-23 cs.CV cs.HC

NeuroSwift: A Lightweight Cross-Subject Framework for fMRI Visual Reconstruction of Complex Scenes

Shiyi Zhang, Dong Liang, Yihang Zhou

机构 * Department of Electronic and Electrical Engineering, Southern University of Science and Technology(电子与电气工程系,南方科技大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院)

Journal ref ACM Multimedia Asia (MMAsia), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09617 2025-10-22 cs.CV

WMamba: Wavelet-based Mamba for Face Forgery Detection

Siran Peng, Tianshuo Zhang, Li Gao, Xiangyu Zhu, Haoyuan Zhang, Kai Pang, Zhen Lei

机构 * CMFT Co., Ltd.(CMFT公司) Guangzhou Pixel Solutions Co., Ltd.(广州像素解决方案公司)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16463 2025-10-21 cs.CV

HGC-Avatar: Hierarchical Gaussian Compression for Streamable Dynamic 3D Avatars

Haocheng Tang, Ruoke Yan, Xinhui Yin, Qi Zhang, Xinfeng Zhang, Siwei Ma, Wen Gao, Chuanmin Jia

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University.(信息处理国家重点实验室,北京大学计算机科学学院) School of Computer Science and Technology, University of Chinese Academy of Sciences.(中国科学院大学计算机科学与技术学院) WICT, State Key Laboratory of Multimedia Information Processing, Peking University.(信息处理国家重点实验室,北京大学WICT) Institute for Clarity in Documentation(文档清晰性研究所) Inria Paris-Rocquencourt(巴黎- Rocquencourt 分部,Inria) Rajiv Gandhi University(拉贾·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(Palmer研究实验室)

Comments ACM International Conference on Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15752 2025-10-20 cs.CV cs.AI

NDM: A Noise-driven Detection and Mitigation Framework against Implicit Sexual Intentions in Text-to-Image Generation

Yitong Sun, Yao Huang, Ruochen Zhang, Huanran Chen, Shouwei Ruan, Ranjie Duan, Xingxing Wei

机构 * Institute of Artificial Intelligence, Beihang University(北航人工智能研究院) College of AI, Tsinghua University(清华人工智能学院) Security Group, Alibaba Group(阿里集团安全组)

Comments 10 pages, 8 figures, accepted by ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12493 2025-10-20 cs.CV

BSGS: Bi-stage 3D Gaussian Splatting for Camera Motion Deblurring

An Zhao, Piaopiao Yu, Zhe Zhu, Mingqiang Wei

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学)

Comments Accept by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13267 2025-10-16 eess.IV cs.HC cs.MM

DIGITWISE: Digital Twin-based Modeling of Adaptive Video Streaming Engagement

Emanuele Artioli, Farzad Tashtarian, Christian Timmerer

Comments ACM Multimedia Systems Conference 2024 (MMSys '24), April 15--18, 2024, Bari, Italy

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16618 2025-10-16 cs.HC cs.CY cs.MM stat.AP

Seeing Isn't Believing: Addressing the Societal Impact of Deepfakes in Low-Tech Environments

Azmine Toushik Wasi, Rahatun Nesa Priti, Mahir Absar Khan, Abdur Rahman, Mst Rafia Islam

Comments Accepted to ACM MM 2025 Workshop Diffusion of Harmful Content on Online Web (DHOW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04247 2025-10-16 cs.IR cs.MM

I$^3$-MRec: Invariant Learning with Information Bottleneck for Incomplete Modality Recommendation

Huilin Chen, Miaomiao Cai, Fan Liu, Zhiyong Cheng, Richang Hong, Meng Wang

Comments ACM Multimedia 2025 Accepted

Journal ref In Proceedings of the 33st ACM International Conference on Multimedia (MM '25), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15298 2025-10-16 cs.CV cs.MM

MEGC2025: Micro-Expression Grand Challenge on Spot Then Recognize and Visual Question Answering

Xinqi Fan, Jingting Li, John See, Moi Hoon Yap, Wen-Huang Cheng, Xiaobai Li, Xiaopeng Hong, Su-Jing Wang, Adrian K. Davision

机构 * Department of Computing and Mathematics, Manchester Metropolitan University(计算与数学系,曼彻斯特 Metropolitan 大学) State Key Laboratory of Cognitive Science and Mental Health, Institute of Psychology, Chinese Academy of Sciences(认知科学与心理健康国家重点实验室,心理学研究所,中国科学院) Department of Psychology, University of the Chinese Academy of Sciences(心理学系,中国科学院大学) National Taiwan University(台湾大学) Zhejiang University(浙江大学) University of Oulu(奥卢大学) Harbin Institute of Technology(哈尔滨工业大学)

Comments Micro-Expression Grand Challenge (MEGC) at ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏