arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

ACM International Conference on Multimedia · 会议 · Multimedia

2025-08-12 至 2025-08-12 共收录 14
2508.07918 2025-08-12 cs.CV

RSVLM-QA: A Benchmark Dataset for Remote Sensing Vision Language Model-based Question Answering

Xing Zi, Jinghao Xiao, Yunxiao Shi, Xian Tao, Jun Li, Ali Braytee, Mukesh Prasad

机构 * School of Computer Science, University of Technology Sydney(技术悉尼大学计算机科学学院) SEDE, University of Technology Sydney(技术悉尼大学SEDE) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

Comments This paper has been accepted to the proceedings of the 33rd ACM International Multimedia Conference (ACM Multimedia 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07766 2025-08-12 cs.CV cs.AI

UniSVG: A Unified Dataset for Vector Graphic Understanding and Generation with Multimodal Large Language Models

Jinke Li, Jiarui Yu, Chenxing Wei, Hande Dong, Qiang Lin, Liangjing Yang, Zhicai Wang, Yanbin Hao

机构 * Zhejiang University(浙江大学) Tencent(腾讯) Shenzhen University(深圳大学) Hefei University of Technology(合肥工业大学)

Comments Accepted at ACM MM 2025 Dataset Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07608 2025-08-12 cs.MM cs.CV cs.SD eess.AS

AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition

Junxiao Xue, Xiaozhen Liu, Xuecheng Wu, Xinyi Yin, Danlei Huang, Fei Yu

机构 * Zhengzhou University(郑州大学) Zhejiang Lab(浙江实验室) Xi'an Jiaotong University(西安交通大学)

Comments Accepted by the ACM MM 2025 Workshop on SVC

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07603 2025-08-12 cs.CV

LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation

Wenhui Song, Hanhui Li, Jiehui Huang, Panwen Hu, Yuhao Cheng, Long Chen, Yiqiang Yan, Xiaodan Liang

机构 * Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) Hong Kong University of Science and Technology(香港科学与技术大学) Mohamed bin Zayed University of Artificial Intelligence(马尔代夫 bin Zayed 人工智能大学) Lenovo Research(联想研究)

Comments Accepted to ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07596 2025-08-12 cs.CV

From Prediction to Explanation: Multimodal, Explainable, and Interactive Deepfake Detection Framework for Non-Expert Users

Shahroz Tariq, Simon S. Woo, Priyanka Singh, Irena Irmalasari, Saakshi Gupta, Dev Gupta

机构 * Sungkyunkwan University, S. Korea(顺天大学) University of Queensland, Australia(昆士兰大学)

Comments 11 pages, 3 tables, 5 figures, accepted for publicaiton in the 33rd ACM International Conference on Multimedia (MM '25), October 27-31, 2025, Dublin, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07337 2025-08-12 eess.AS cs.CV

KLASSify to Verify: Audio-Visual Deepfake Detection Using SSL-based Audio and Handcrafted Visual Features

Ivan Kukanov, Jun Wah Ng

机构 * KLASS Engineering and Solutions Singapore(KLASS工程与解决方案新加坡)

Comments 7 pages, accepted to the 33rd ACM International Conference on Multimedia (MM'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05658 2025-08-12 cs.CR cs.CV cs.MM

Universally Unfiltered and Unseen:Input-Agnostic Multimodal Jailbreaks against Text-to-Image Model Safeguards

Song Yan, Hui Wei, Jinlong Fei, Guoliang Yang, Zhengyu Zhao, Zheng Wang

机构 * Information Engineering University Zhengzhou China School of Computer Science, \ University Wuhan China Xi’an Jiaotong University Xi’an China Wuhan University Wuhan China Information Engineering University School of Computer Science, \ University Xi’an Jiaotong University Wuhan University

Comments This paper has been accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05016 2025-08-12 cs.CV eess.IV

AU-IQA: A Benchmark Dataset for Perceptual Quality Assessment of AI-Enhanced User-Generated Content

Shushi Wang, Chunyi Li, Zicheng Zhang, Han Zhou, Wei Dong, Jun Chen, Guangtao Zhai, Xiaohong Liu

机构 * Shanghai Jiao Tong University(上海交通大学) McMaster University(麦斯特大学) Suzhou Key Laboratory of Artificial Intelligence(苏州人工智能重点实验室)

Comments Accepted by ACMMM 2025 Datasets Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07526 2025-08-12 cs.SD eess.AS

DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram Reconstruction

Cunhang Fan, Sheng Zhang, Jingjing Zhang, Enrui Liu, Xinhui Li, Gangming Zhao, Zhao Lv

机构 * School of Computer Science and Technology(计算机科学与技术学院) Anhui University(安徽大学)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04377 2025-08-12 cs.CV cs.CL cs.MM

Multi-Modal Semantic Parsing for the Interpretation of Tombstone Inscriptions

Xiao Zhang, Johan Bos

机构 * University of Groningen(格罗宁根大学)

Comments ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00425 2025-08-12 cs.CV cs.AI

MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization

JiangYong Yu, Sifan Zhou, Dawei Yang, Shuo Wang, Shuoyu Li, Xing Hu, Chen Xu, Zukang Xu, Changyong Shu, Zhihang Yuan

机构 * Southeast University(东南大学) Xi'an Jiaotong University(西安交通大学)

Comments Accepted by ACM MM 2025. First PTQ solution for Multimodal large language models applicable to 5 mainstream MLLMs

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17632 2025-08-12 cs.AI cs.CV cs.MM

D-Judge: How Far Are We? Assessing the Discrepancies Between AI-synthesized and Natural Images through Multimodal Guidance

Renyang Liu, Ziyu Lyu, Wei Zhou, See-Kiong Ng

机构 * School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University(中山大学信息科学与技术学院(深圳校区)) Institute of Data Science, National University of Singapore(新加坡国立大学数据科学研究所) College of Modern Engineering and the Engineering Research Center of Cyberspace, Yunnan University(云南大学现代工程学院及空天信息工程研究中心)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06917 2025-08-12 q-bio.QM cs.AI

CROP: Integrating Topological and Spatial Structures via Cross-View Prefixes for Molecular LLMs

Jianting Tang, Yubo Wang, Haoyu Cao, Linli Xu

机构 * University of Science and Technology of China(科学技术大学) State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室)

Comments Accepted to ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13623 2025-08-12 cs.CV

Prompt-Softbox-Prompt: A Free-Text Embedding Control for Image Editing

Yitong Yang, Yinglin Wang, Tian Zhang, Jing Wang, Shuting He

机构 * Shanghai University of Finance and Economics(上海财经大学) School of Computing and Artificial Intelligence(计算机与人工智能学院) MoE Key Laboratory of Interdisciplinary Research of Computation and Economics(交叉计算与经济学联合实验室)

Comments ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏