arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

ACM International Conference on Multimedia · 会议 · Multimedia

共收录 1944
2508.08754 2025-08-13 cs.GR cs.CV cs.MM

Exploring Palette based Color Guidance in Diffusion Models

Qianru Qiu, Jiafeng Mao, Xueting Wang

机构 * CyberAgent Inc.(CyberAgent公司)

Comments Accepted to ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07918 2025-08-12 cs.CV

RSVLM-QA: A Benchmark Dataset for Remote Sensing Vision Language Model-based Question Answering

Xing Zi, Jinghao Xiao, Yunxiao Shi, Xian Tao, Jun Li, Ali Braytee, Mukesh Prasad

机构 * School of Computer Science, University of Technology Sydney(技术悉尼大学计算机科学学院) SEDE, University of Technology Sydney(技术悉尼大学SEDE) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

Comments This paper has been accepted to the proceedings of the 33rd ACM International Multimedia Conference (ACM Multimedia 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07766 2025-08-12 cs.CV cs.AI

UniSVG: A Unified Dataset for Vector Graphic Understanding and Generation with Multimodal Large Language Models

Jinke Li, Jiarui Yu, Chenxing Wei, Hande Dong, Qiang Lin, Liangjing Yang, Zhicai Wang, Yanbin Hao

机构 * Zhejiang University(浙江大学) Tencent(腾讯) Shenzhen University(深圳大学) Hefei University of Technology(合肥工业大学)

Comments Accepted at ACM MM 2025 Dataset Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07608 2025-08-12 cs.MM cs.CV cs.SD eess.AS

AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition

Junxiao Xue, Xiaozhen Liu, Xuecheng Wu, Xinyi Yin, Danlei Huang, Fei Yu

机构 * Zhengzhou University(郑州大学) Zhejiang Lab(浙江实验室) Xi'an Jiaotong University(西安交通大学)

Comments Accepted by the ACM MM 2025 Workshop on SVC

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07603 2025-08-12 cs.CV

LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation

Wenhui Song, Hanhui Li, Jiehui Huang, Panwen Hu, Yuhao Cheng, Long Chen, Yiqiang Yan, Xiaodan Liang

机构 * Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) Hong Kong University of Science and Technology(香港科学与技术大学) Mohamed bin Zayed University of Artificial Intelligence(马尔代夫 bin Zayed 人工智能大学) Lenovo Research(联想研究)

Comments Accepted to ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07596 2025-08-12 cs.CV

From Prediction to Explanation: Multimodal, Explainable, and Interactive Deepfake Detection Framework for Non-Expert Users

Shahroz Tariq, Simon S. Woo, Priyanka Singh, Irena Irmalasari, Saakshi Gupta, Dev Gupta

机构 * Sungkyunkwan University, S. Korea(顺天大学) University of Queensland, Australia(昆士兰大学)

Comments 11 pages, 3 tables, 5 figures, accepted for publicaiton in the 33rd ACM International Conference on Multimedia (MM '25), October 27-31, 2025, Dublin, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07337 2025-08-12 eess.AS cs.CV

KLASSify to Verify: Audio-Visual Deepfake Detection Using SSL-based Audio and Handcrafted Visual Features

Ivan Kukanov, Jun Wah Ng

机构 * KLASS Engineering and Solutions Singapore(KLASS工程与解决方案新加坡)

Comments 7 pages, accepted to the 33rd ACM International Conference on Multimedia (MM'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05658 2025-08-12 cs.CR cs.CV cs.MM

Universally Unfiltered and Unseen:Input-Agnostic Multimodal Jailbreaks against Text-to-Image Model Safeguards

Song Yan, Hui Wei, Jinlong Fei, Guoliang Yang, Zhengyu Zhao, Zheng Wang

机构 * Information Engineering University Zhengzhou China School of Computer Science, \ University Wuhan China Xi’an Jiaotong University Xi’an China Wuhan University Wuhan China Information Engineering University School of Computer Science, \ University Xi’an Jiaotong University Wuhan University

Comments This paper has been accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05016 2025-08-12 cs.CV eess.IV

AU-IQA: A Benchmark Dataset for Perceptual Quality Assessment of AI-Enhanced User-Generated Content

Shushi Wang, Chunyi Li, Zicheng Zhang, Han Zhou, Wei Dong, Jun Chen, Guangtao Zhai, Xiaohong Liu

机构 * Shanghai Jiao Tong University(上海交通大学) McMaster University(麦斯特大学) Suzhou Key Laboratory of Artificial Intelligence(苏州人工智能重点实验室)

Comments Accepted by ACMMM 2025 Datasets Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07526 2025-08-12 cs.SD eess.AS

DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram Reconstruction

Cunhang Fan, Sheng Zhang, Jingjing Zhang, Enrui Liu, Xinhui Li, Gangming Zhao, Zhao Lv

机构 * School of Computer Science and Technology(计算机科学与技术学院) Anhui University(安徽大学)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04377 2025-08-12 cs.CV cs.CL cs.MM

Multi-Modal Semantic Parsing for the Interpretation of Tombstone Inscriptions

Xiao Zhang, Johan Bos

机构 * University of Groningen(格罗宁根大学)

Comments ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00425 2025-08-12 cs.CV cs.AI

MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization

JiangYong Yu, Sifan Zhou, Dawei Yang, Shuo Wang, Shuoyu Li, Xing Hu, Chen Xu, Zukang Xu, Changyong Shu, Zhihang Yuan

机构 * Southeast University(东南大学) Xi'an Jiaotong University(西安交通大学)

Comments Accepted by ACM MM 2025. First PTQ solution for Multimodal large language models applicable to 5 mainstream MLLMs

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17632 2025-08-12 cs.AI cs.CV cs.MM

D-Judge: How Far Are We? Assessing the Discrepancies Between AI-synthesized and Natural Images through Multimodal Guidance

Renyang Liu, Ziyu Lyu, Wei Zhou, See-Kiong Ng

机构 * School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University(中山大学信息科学与技术学院(深圳校区)) Institute of Data Science, National University of Singapore(新加坡国立大学数据科学研究所) College of Modern Engineering and the Engineering Research Center of Cyberspace, Yunnan University(云南大学现代工程学院及空天信息工程研究中心)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06917 2025-08-12 q-bio.QM cs.AI

CROP: Integrating Topological and Spatial Structures via Cross-View Prefixes for Molecular LLMs

Jianting Tang, Yubo Wang, Haoyu Cao, Linli Xu

机构 * University of Science and Technology of China(科学技术大学) State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室)

Comments Accepted to ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13623 2025-08-12 cs.CV

Prompt-Softbox-Prompt: A Free-Text Embedding Control for Image Editing

Yitong Yang, Yinglin Wang, Tian Zhang, Jing Wang, Shuting He

机构 * Shanghai University of Finance and Economics(上海财经大学) School of Computing and Artificial Intelligence(计算机与人工智能学院) MoE Key Laboratory of Interdisciplinary Research of Computation and Economics(交叉计算与经济学联合实验室)

Comments ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06382 2025-08-11 cs.CV

Text as Any-Modality for Zero-Shot Classification by Consistent Prompt Tuning

Xiangyu Wu, Feng Yu, Yang Yang, Jianfeng Lu

机构 * Nanjing University of Science

Comments Accepted for publication at ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05741 2025-08-11 cs.CV

Dome-DETR: DETR with Density-Oriented Feature-Query Manipulation for Efficient Tiny Object Detection

Zhangchi Hu, Peixi Wu, Jie Chen, Huyue Zhu, Yijun Wang, Yansong Peng, Hebei Li, Xiaoyan Sun

机构 * University of Science and Technology of China(科学技术大学) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(人工智能研究所)

Comments Accepted by ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05585 2025-08-08 cs.CV

DART: Dual Adaptive Refinement Transfer for Open-Vocabulary Multi-Label Recognition

Haijing Liu, Tao Pu, Hefeng Wu, Keze Wang, Liang Lin

机构 * Sun Yat-sen University(中山大学)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05343 2025-08-08 cs.CV

3DGabSplat: 3D Gabor Splatting for Frequency-adaptive Radiance Field Rendering

Junyu Zhou, Yuyang Huang, Wenrui Dai, Junni Zou, Ziyang Zheng, Nuowen Kan, Chenglin Li, Hongkai Xiong

机构 * Shanghai Jiao Tong University(上海交通大学)

Comments Accepted by ACM MM'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05087 2025-08-08 cs.MM cs.AI cs.CL cs.CR

JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual Steering

Renmiao Chen, Shiyao Cui, Xuancheng Huang, Chengwei Pan, Victor Shea-Jay Huang, QingLin Zhang, Xuan Ouyang, Zhexin Zhang, Hongning Wang, Minlie Huang

机构 * CoAI group, DCST, Tsinghua University(清华大学DCST学院) Beihang University(北航大学)

Comments 10 pages, 3 tables, 2 figures, to appear in the Proceedings of the 33rd ACM International Conference on Multimedia (MM '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04942 2025-08-08 cs.CV

Accelerating Conditional Prompt Learning via Masked Image Modeling for Vision-Language Models

Phuoc-Nguyen Bui, Khanh-Binh Nguyen, Hyunseung Choo

机构 * Convergence Research Institute, Sungkyunkwan University(convergence research institute, 首尔大学) School of Information Technology, Deakin University(信息科技学院, 德金大学) Department of Electrical and Computer Engineering, Sungkyunkwan University(电气与计算机工程系, 首尔大学)

Comments ACMMM-LAVA 2025, 10 pages, camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04723 2025-08-08 cs.SD cs.AI eess.AS

Wearable Music2Emotion : Assessing Emotions Induced by AI-Generated Music through Portable EEG-fNIRS Fusion

Sha Zhao, Song Yi, Yangxuan Zhou, Jiadong Pan, Jiquan Wang, Jie Xia, Shijian Li, Shurong Dong, Gang Pan

机构 * Zhejiang University(浙江大学) Hangzhou RongNao Technology Co., Ltd(杭州融脑科技有限公司)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04585 2025-08-08 eess.AS

UniTalker: Conversational Speech-Visual Synthesis

Yifan Hu, Rui Liu, Yi Ren, Xiang Yin, Haizhou Li

Comments 15 pages, 8 figures, Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04197 2025-08-07 cs.CV cs.AI

Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective

Yan Zhang, Gangyan Zeng, Daiqing Wu, Huawen Shen, Binbin Li, Yu Zhou, Can Ma, Xiaojun Bi

机构 * Institute of Information Engineering, Chinese Academy of Sciences School of Cyber Security, University of Chinese Academy of Sciences Beijing China School of Cyber Science Engineering, Nanjing University of Science VCIP \& TMCC \& DISSec, College of Computer Science, Nankai University Tianjin China Key Laboratory of Ethnic Language Intelligent Analysis Security Governance of MOE, Minzu University of China Beijing China Institute of Information Engineering, Chinese Academy of Sciences School of Cyber Security, University of Chinese Academy of Sciences VCIP \& TMCC \& DISSec, College of Computer Science, Nankai University Security Governance of MOE, Minzu University of China

Comments Accepted by 2025 ACM MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04129 2025-08-07 cs.CV

SVC 2025: the First Multimodal Deception Detection Challenge

Xun Lin, Xiaobao Guo, Taorui Wang, Yingjie Ma, Jiajian Huang, Jiayu Zhang, Junzhe Cao, Zitong Yu

机构 * Great Bay University(大亚湾大学) Nanyang Technological University(南洋理工大学)

Comments Accepted by Workshop SVC of ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04050 2025-08-07 cs.CV

DOMR: Establishing Cross-View Segmentation via Dense Object Matching

Jitong Liao, Yulu Gao, Shaofei Huang, Jialin Gao, Jie Lei, Ronghua Liang, Si Liu

机构 * Hangzhou International Innovation Institute, Beihang University(北京航空航天大学杭州国际创新研究院) Faculty of Science and Technology, University of Macau(澳门大学科学与技术学院) College of Computer Science and Technology, Zhejiang University of Technology(浙江工业大学计算机科学与技术学院) School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10576 2025-08-07 cs.GR

Robust Photo-Realistic Hand Gesture Generation: from Single View to Multiple View

Qifan Fu, Xu Chen, Muhammad Asad, Shanxin Yuan, Changjae Oh, Gregory Slabaugh

Comments This nine pages paper has been accepted for publication in Proceedings of the 33rd ACM International Conference on Multimedia (ACM MM 2025). This is the author's version which has not been fully edited and content may change prior to final publication. Citation information: DOI https://doi.org/10.1145/3746027.3755828

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18087 2025-08-07 cs.CV

Disentangle Identity, Cooperate Emotion: Correlation-Aware Emotional Talking Portrait Generation

Weipeng Tan, Chuming Lin, Chengming Xu, FeiFan Xu, Xiaobin Hu, Xiaozhong Ji, Junwei Zhu, Chengjie Wang, Yanwei Fu

机构 * Fudan University(复旦大学) Tencent, YouTu Lab(腾讯、YouTu实验室)

Comments Accepted by ACM MM'25. arXiv admin note: text overlap with arXiv:2409.03270

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13053 2025-08-07 cs.CL

Evaluating the Robustness of Multimodal Agents Against Active Environmental Injection Attacks

Yurun Chen, Xavier Hu, Keting Yin, Juncheng Li, Shengyu Zhang

机构 * School of Software Technology Zhejiang University Hangzhou Zhejiang China(软件技术学院浙江大学杭州浙江中国) Zhejiang University(浙江大学)

Comments Accepted at ACM MM 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10352 2025-08-06 eess.AS cs.CL

Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis

Yifan Yang, Shujie Liu, Jinyu Li, Yuxuan Hu, Haibin Wu, Hui Wang, Jianwei Yu, Lingwei Meng, Haiyang Sun, Yanqing Liu, Yan Lu, Kai Yu, Xie Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Microsoft Corporation(微软公司) Shanghai Jiao Tong University, Shanghai Innovation Institute(上海交通大学上海创新研究院)

Comments Accepted in ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏