arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

ACM International Conference on Multimedia · 会议 · Multimedia

共收录 1940
2510.12267 2025-10-15 cs.CV

SpineBench: Benchmarking Multimodal LLMs for Spinal Pathology Analysis

Chenghanyu Zhang, Zekun Li, Peipei Li, Xing Cui, Shuhan Xia, Weixiang Yan, Yiqiao Zhang, Qianyu Zhuang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) University of California, Santa Barbara(加州大学圣巴巴拉分校) Peking Union Medical College Hospital(北京佑安医学大学医院)

Comments Proceedings of the 33rd ACM International Conference on Multimedia,ACMMM 2025 Dataset Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12056 2025-10-15 cs.CV

APGNet: Adaptive Prior-Guided for Underwater Camouflaged Object Detection

Xinxin Huang, Han Sun, Junmin Cai, Ningzhong Liu, Huiyu Zhou

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学) University of Leicester(莱斯特大学)

Comments 6 pages. accepted by ACM MM Asia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11738 2025-10-15 cs.SD cs.AI cs.CV cs.MM

SeeingSounds: Learning Audio-to-Visual Alignment via Text

Simone Carnemolla, Matteo Pennisi, Chiara Russo, Simone Palazzo, Daniela Giordano, Concetto Spampinato

机构 * University of Catania(卡塔尼亚大学)

Comments accepted to ACM Multimedia Asia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01427 2025-10-15 cs.CV cs.AI

Capturing More: Learning Multi-Domain Representations for Robust Online Handwriting Verification

Peirong Zhang, Kai Ding, Lianwen Jin

机构 * South China University of Technology(华南理工大学) INTSIG Information Co. Ltd(INTSIG信息有限公司) INTSIG-SCUT Joint Lab on Document Analysis and Recognition(INTSIG-SCUT文档分析与识别联合实验室) SCUT-Zhuhai Institute of Modern Industrial Innovation(华南理工大学珠海现代工业创新研究院)

Comments Accepted to ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05970 2025-10-14 cs.CV

Automatic Synthesis of High-Quality Triplet Data for Composed Image Retrieval

Haiwen Li, Delong Liu, Zhaohui Hou, Zhicheng Zhao, Fei Su

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

Comments This paper was originally submitted to ACM MM 2025 on April 12, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09479 2025-10-14 cs.AI cs.CL

Draw with Thought: Unleashing Multimodal Reasoning for Scientific Diagram Generation

Zhiqing Cui, Jiahao Yuan, Hanqing Wang, Yanshu Li, Chenxu Du, Zhenglong Ding

机构 * Nanjing University of Information Science \& Technology Nanjing China East China Normal University Shanghai China The Hong Kong University of Science Brown University Providence America Southwest Jiaotong University Chengdu China Nanjing University of Information Science \& Technology East China Normal University Brown University Southwest Jiaotong University

Comments 10 pages, 5 figures, accepted to appear in the Proceedings of the 33rd ACM International Conference on Multimedia (MM '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17360 2025-10-14 cs.CV

UniRGB-IR: A Unified Framework for Visible-Infrared Semantic Tasks via Adapter Tuning

Maoxun Yuan, Bo Cui, Tianyi Zhao, Jiayi Wang, Shan Fu, Xue Yang, Xingxing Wei

机构 * Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院) School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) CTTL-Terminal, China Academy of Information and Communications Technology(信息通信技术中国科学院CTTL终端) School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(上海交通大学自动化与智能感知学院)

Comments 10 pages, 6 figures, Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10022 2025-10-14 cs.CV

Q-Adapter: Visual Query Adapter for Extracting Textually-related Features in Video Captioning

Junan Chen, Trung Thanh Nguyen, Takahiro Komamizu, Ichiro Ide

机构 * Nagoya University(名古屋大学)

Comments ACM Multimedia Asia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01181 2025-10-14 cs.AI cs.CV cs.MM cs.SD eess.AS

Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion Reasoning

Zhiyuan Han, Beier Zhu, Yanlong Xu, Peipei Song, Xun Yang

机构 * University of Science and Technology of China(科学技术大学) Nanyang Technological University(南洋理工大学)

Comments ACM Multimedia 2025 Oral Code: https://github.com/ZhiyuanHan-Aaron/MoSEAR Project Page: https://zhiyuanhan-aaron.github.io/MoSEAR-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08004 2025-10-10 cs.SD cs.MM eess.AS

Personality-Enhanced Multimodal Depression Detection in the Elderly

Honghong Wang, Jing Deng, Rong Zheng

机构 * Beijing Fosafer Information Technology Co., Ltd.(北京福萨弗信息科技有限公司)

Comments 6 pages,2 figures,accepted by ACM Multimedia Asia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02270 2025-10-09 cs.CV cs.AI

Sustainable Self-evolution Adversarial Training

Wenxuan Wang, Chenglei Wang, Huihui Qi, Menghao Ye, Xuelin Qian, Peng Wang, Yanning Zhang

机构 * School of Computer Science, Northwestern Polytechnical University,National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology(计算机科学学院,西北工业大学,集成空天地海大数据应用技术国家工程实验室) School of Automation, Northwestern Polytechnical University,the BRain and Artificial INtelligence Lab(自动化学院,西北工业大学,脑与人工智能实验室)

Comments Accepted to ACMMM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05984 2025-10-08 cs.SD cs.AI eess.AS

ECTSpeech: Enhancing Efficient Speech Synthesis via Easy Consistency Tuning

Tao Zhu, Yinfeng Yu, Liejun Wang, Fuchun Sun, Wendong Zheng

机构 * School of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院) Joint International Research Laboratory of Silk Road Multilingual Cognitive Computing(丝绸之路多语种认知计算联合国际实验室) Tsinghua University(清华大学) Tianjin University of Technology(天津工业大学)

Comments Accepted for publication by Proceedings of the 2025 ACM Multimedia Asia Conference(MMAsia '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04712 2025-10-07 cs.CV cs.HC cs.MM

ReactDiff: Fundamental Multiple Appropriate Facial Reaction Diffusion Model

Luo Cheng, Song Siyang, Yan Siyuan, Yu Zhen, Ge Zongyuan

机构 * Monash University(墨尔本大学) University of Exeter(埃克塞特大学)

Comments Accepted to ACM Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03840 2025-10-07 cs.CV

Mirage: Unveiling Hidden Artifacts in Synthetic Images with Large Vision-Language Models

Pranav Sharma, Shivank Garg, Durga Toshniwal

机构 * Indian Institute of Technology Roorkee(印度理工学院罗奥里分校)

Comments ACM MM'25, MALLM Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09887 2025-10-07 cs.LG cs.AI eess.SP

TolerantECG: A Foundation Model for Imperfect Electrocardiogram

Huynh Dang Nguyen, Trong-Thang Pham, Ngan Le, Van Nguyen

机构 * AI Center, FPT Software(FPT软件人工智能中心) EECS, University of Arkansas(亚利桑那大学电子工程与计算机科学系)

Comments Accepted at ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09747 2025-10-07 cs.NE

BrainFLORA: Uncovering Brain Concept Representation via Multimodal Neural Embeddings

Dongyang Li, Haoyang Qin, Mingyang Wu, Chen Wei, Quanying Liu

Comments ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09255 2025-10-07 cs.CV

FVQ: A Large-Scale Dataset and an LMM-based Method for Face Video Quality Assessment

Sijing Wu, Yunhao Li, Ziwen Xu, Yixuan Gao, Huiyu Duan, Wei Sun, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) East China Normal University(华东师范大学)

Comments Accepted by ACM MM 2025. Project page: https://github.com/wsj-sjtu/FVQ

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03038 2025-10-06 cs.LG cs.AI cs.IR

CHORD: Customizing Hybrid-precision On-device Model for Sequential Recommendation with Device-cloud Collaboration

Tianqi Liu, Kairui Fu, Shengyu Zhang, Wenyan Fan, Zhaocheng Du, Jieming Zhu, Fan Wu, Fei Wu

机构 * Zhejiang University(浙江大学) Shanghai Institute for Advanced Study of Zhejiang University(浙江大学上海高等研究院) Huawei Noah’s Ark Lab(华为诺亚实验室) Shanghai Jiao Tong University(上海交通大学)

Comments accepted by ACM MM'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02787 2025-10-06 cs.CV

OTR: Synthesizing Overlay Text Dataset for Text Removal

Jan Zdenek, Wataru Shimoda, Kota Yamaguchi

机构 * CyberAgent Tokyo Japan(CyberAgent东京日本)

Comments This is the author's version of the work. It is posted here for your personal use. Not for redistribution. The definitive Version of Record was published in Proceedings of the 33rd ACM International Conference on Multimedia (MM '25), October 27-31, 2025, Dublin, Ireland, https://doi.org/10.1145/3746027.3758297

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20629 2025-10-06 cs.CV cs.AI cs.MM

AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation

Jeongsoo Choi, Ji-Hoon Kim, Kim Sung-Bin, Tae-Hyun Oh, Joon Son Chung

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院) Pohang University of Science and Technology(釜山科学技术大学)

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01585 2025-10-03 cs.CL cs.NI

ReSSFormer: A Recursive Sparse Structured Transformer for Scalable and Long-Context Reasoning

Haochen You, Baojing Liu

机构 * Columbia University(哥伦比亚大学) Hebei Institute of Communications(河北通信学院)

Comments Accepted as a short paper at ACM Multimedia Asia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01578 2025-10-03 cs.LG

Gradient Shaping Beyond Clipping: A Functional Perspective on Update Magnitude Control

Haochen You, Baojing Liu

机构 * Columbia University(哥伦比亚大学) Hebei Institute of Communications(河北通信学院)

Comments Accepted as a conference paper at ACM Multimedia Asia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11997 2025-10-03 cs.LG cs.AI

Can LLMs Find Fraudsters? Multi-level LLM Enhanced Graph Fraud Detection

Tairan Huang, Yili Wang, Qiutong Li, Changlong He, Jianliang Gao

机构 * Central South University(中南大学) Hongkong University of Science and Technology (Guangzhou)(香港科技大学(广州))

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00639 2025-10-02 cs.SD

Reference-free automatic speech severity evaluation using acoustic unit language modelling

Bence Mark Halpern, Tomoki Toda

机构 * Nagoya University(名古屋大学)

Comments 5 pages. Proceedings of the 6th ACM International Conference on Multimedia in Asia Workshops

Journal ref In Proceedings of the 6th ACM International Conference on Multimedia in Asia Workshops (pp. 1-5) (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00647 2025-10-02 cs.CL

MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation

Jinlan Fu, Shenzhen Huangfu, Hao Fei, Yichong Huang, Xiaoyu Shen, Xipeng Qiu, See-Kiong Ng

机构 * National University of Singapore(新加坡国立大学) Fudan University(复旦大学) Harbin Institute of Technology(哈尔滨工业大学) Eastern Institute of Technology(东方技术研究所)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20715 2025-10-02 cs.CV cs.AI

Beyond the Individual: Introducing Group Intention Forecasting with SHOT Dataset

Ruixu Zhang, Yuran Wang, Xinyi Hu, Chaoyu Mai, Wenxuan Liu, Danni Xu, Xian Zhong, Zheng Wang

机构 * National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, School of Computer Science(国家多媒体软件工程研究中心、人工智能研究所、计算机科学学院) Tsinghua University(清华大学) School of Mathematical Sciences, Peking University(北京大学数学科学学院) National Engineering Research Center for Multimedia Software, School of Computer Science(国家多媒体软件工程研究中心、计算机科学学院) School of Computer Science, Peking University(北京大学计算机科学学院) State Key Laboratory for Multimedia Information Processing(多媒体信息处理国家重点实验室) School of Computing, National University of Singapore(新加坡国立大学计算机学院) Hubei Key Laboratory of Transportation Internet of Things, School of Computer Science and Artificial Intelligence(湖北交通物联网重点实验室、计算机科学与人工智能学院) Wuhan University of Technology(武汉理工大学)

Comments ACMMM 2025 Datasets Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26008 2025-10-01 cs.CV cs.AI cs.CG

PFDepth: Heterogeneous Pinhole-Fisheye Joint Depth Estimation via Distortion-aware Gaussian-Splatted Volumetric Fusion

Zhiwei Zhang, Ruikai Xu, Weijian Zhang, Zhizhong Zhang, Xin Tan, Jingyu Gong, Yuan Xie, Lizhuang Ma

机构 * Shanghai Jiao Tong University(上海交通大学) East China Normal University(华东师范大学) Shanghai Key Laboratory of Computer Software Evaluating(上海计算机软件评测测试重点实验室)

Comments Accepted by ACM MM 2025 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18911 2025-09-30 cs.CV

Synthetic-to-Real Camouflaged Object Detection

Zhihao Luo, Luojun Lin, Zheng Lin

机构 * Fuzhou University(福州大学) Tsinghua University(清华大学)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19178 2025-09-30 cs.CV cs.AI cs.CL cs.IR cs.LG

Reversed in Time: A Novel Temporal-Emphasized Benchmark for Cross-Modal Video-Text Retrieval

Yang Du, Yuqi Liu, Qin Jin

机构 * Renmin University of China(中国人民大学)

Comments ACMMM 2024 poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22168 2025-09-29 cs.HC cs.AI

Teaching AI to Feel: A Collaborative, Full-Body Exploration of Emotive Communication

Esen K. Tütüncü, Lissette Lemus, Kris Pilcher, Holger Sprengel, Jordi Sabater-Mir

机构 * Institute of Neurosciences of the University of Barcelona Spain Artificial Intelligence Research Institute (IIIA-CSIC) Barcelona Spain Massachusetts Institute of Technology Cambridge USA ESPRONCEDA Institute of Art \& Culture Barcelona Spain Institute of Neurosciences of the University of Artificial Intelligence Research Institute (IIIA-CSIC) Massachusetts Institute of Technology ESPRONCEDA Institute of Art \& Culture

Comments 9 pages, 10 Figures, ACM MM'25

详情

展开后加载摘要…

URL PDF HTML 收藏