arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

ACM International Conference on Multimedia · 会议 · Multimedia

共收录 1944
2507.09647 2025-07-18 cs.MM cs.AI

KEN: Knowledge Augmentation and Emotion Guidance Network for Multimodal Fake News Detection

Peican Zhu, Yubo Jing, Le Cheng, Keke Tang, Yangming Guo

机构 * Northwestern Polytechnical University(西北工业大学) Guangzhou University(广东大学)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12461 2025-07-17 cs.CV cs.AI

Interpreting Radiologist's Intention from Eye Movements in Chest X-ray Diagnosis

Trong-Thang Pham, Anh Nguyen, Zhigang Deng, Carol C. Wu, Hien Van Nguyen, Ngan Le

机构 * University of Arkansas(亚拉巴马大学) University of Liverpool(利物浦大学) University of Houston(休斯顿大学) MD Anderson Cancer Center(MD安德森癌症中心)

Comments ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12062 2025-07-17 cs.CV

MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic Learning

Hongxu Ma, Guanshuo Wang, Fufu Yu, Qiong Jia, Shouhong Ding

机构 * Fudan University(复旦大学) Tencent Youtu Lab(腾讯优图实验室)

Comments Accepted by ACM MM'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11661 2025-07-17 cs.CL cs.AI

Partitioner Guided Modal Learning Framework

Guimin Hu, Yi Xin, Lijie Hu, Zhihong Zhu, Hasti Seifi

机构 * Guangdong University of Technology(广东工业大学) University of Copenhagen(哥本哈根大学) Nanjing University(南京大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Tencent(腾讯) Arizona State University(亚利桑那州立大学)

Comments acm multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10432 2025-07-17 cs.CV

Text-Visual Semantic Constrained AI-Generated Image Quality Assessment

Qiang Li, Qingsen Yan, Haojian Huang, Peng Wu, Haokui Zhang, Yanning Zhang

机构 * Northwestern Polytechnical University(西北工业大学) The University of Hong Kong(香港大学)

Comments 9 pages, 5 figures, Accepted at ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10358 2025-07-15 cs.CV

Fine-Grained Zero-Shot Object Detection

Hongxu Ma, Chenbo Zhang, Lu Zhang, Jiaogen Zhou, Jihong Guan, Shuigeng Zhou

机构 * Fudan University(复旦大学) School of Geography and Planning, Huaiyin Normal University(地理与规划学院,淮阴师范学院) Tongji University(同济大学)

Comments Accepted by ACM MM'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09876 2025-07-15 cs.CV cs.AI cs.CL

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models

Yongheng Zhang, Xu Liu, Ruihan Tao, Qiguang Chen, Hao Fei, Wanxiang Che, Libo Qin

机构 * School of Computer Science and Engineering, Central South University(计算机科学与工程学院,中南大学) Research Center for SCIR, Harbin Institute of Technology(SCIR研究中心,哈尔滨工业大学) NExT Research Center, National University of Singapore(NExT研究中心,新加坡国立大学)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04840 2025-07-15 cs.MM

Unsupervised Ego- and Exo-centric Dense Procedural Activity Captioning via Gaze Consensus Adaptation

Zhaofeng Shi, Heqian Qiu, Lanxiao Wang, Qingbo Wu, Fanman Meng, Hongliang Li

Comments ACM International Conference on Multimedia(ACM MM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06678 2025-07-15 cs.CV

Gamma: Toward Generic Image Assessment with Mixture of Assessment Experts

Hantao Zhou, Rui Yang, Longxiang Tang, Guanyi Qin, Runze Hu, Xiu Li

机构 * Tsinghua University(清华大学) The University of Hong Kong(香港大学) National University of Singapore(新加坡国立大学) Beijing Institute of Technology(北京理工大学)

Comments Accepted to ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09334 2025-07-15 cs.CV

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding

Wencan Huang, Daizong Liu, Wei Hu

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07902 2025-07-11 cs.CV

MIRA: A Novel Framework for Fusing Modalities in Medical RAG

Jinhong Wang, Tajamul Ashraf, Zongyan Han, Jorma Laaksonen, Rao Mohammad Anwer

机构 * Department of Computer Vision, MBZUAI(视觉计算系,MBZUAI) Department of Computer Science, Aalto University(计算机科学系,阿alto大学)

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07708 2025-07-11 cs.CV

Motion-Aware Adaptive Pixel Pruning for Efficient Local Motion Deblurring

Wei Shang, Dongwei Ren, Wanying Zhang, Pengfei Zhu, Qinghua Hu, Wangmeng Zuo

机构 * Harbin Institute of Technology \& City University of Hong Kong School of Computer Science Technology, Harbin Institute of Technology Tianjin University \& Low-Altitude Intelligence Lab, Xiong'an National Innovation Center Technology Co., Ltd College of Intelligence Computing, Tianjin University

Comments Accepted by ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07015 2025-07-10 cs.CV cs.LG cs.MM

MST-Distill: Mixture of Specialized Teachers for Cross-Modal Knowledge Distillation

Hui Li, Pengfei Yang, Juanyang Chen, Le Dong, Yanxin Chen, Quan Wang

机构 * Xidian University(西安电子科技大学)

Comments Accepted to ACM MM 2025 (The 33rd ACM International Conference on Multimedia)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06671 2025-07-10 cs.CV

FlexGaussian: Flexible and Cost-Effective Training-Free Compression for 3D Gaussian Splatting

Boyuan Tian, Qizhe Gao, Siran Xianyu, Xiaotong Cui, Minjia Zhang

Comments To appear at ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14307 2025-07-10 cs.CV cs.AI

DilateQuant: Accurate and Efficient Diffusion Quantization via Weight Dilation

Xuewen Liu, Zhikai Li, Minhao Jiang, Mengjuan Chen, Jianquan Li, Qingyi Gu

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

Comments ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05939 2025-07-09 cs.CL cs.MM

Remember Past, Anticipate Future: Learning Continual Multimodal Misinformation Detectors

Bing Wang, Ximing Li, Mengzhe Ye, Changchun Li, Bo Fu, Jianfeng Qu, Lin Yuanbo Wu

机构 * College of Computer Science and Technology, Jilin University(吉林大学计算机科学与技术学院) College of Software, Jilin University(吉林大学软件学院) School of Computer and Artificial Intelligence, Liaoning Normal University(辽宁师范大学计算机与人工智能学院) School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院) Department of Computer Science, Swansea University(斯旺西大学计算机科学系)

Comments Accepted by ACM MM 2025. 10 pages, 6 figures. Code: https://github.com/wangbing1416/DAEDCMD

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11079 2025-07-09 cs.SD cs.CL eess.AS

ALLM4ADD: Unlocking the Capabilities of Audio Large Language Models for Audio Deepfake Detection

Hao Gu, Jiangyan Yi, Chenglong Wang, Jianhua Tao, Zheng Lian, Jiayi He, Yong Ren, Yujie Chen, Zhengqi Wen

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Department of Automation, Tsinghua University(清华大学自动化系) Taizhou University(台州大学) Anhui University(安徽大学) Beijing National Research Center for Information Science and Technology,Tsinghua University(北京信息科学与技术国家研究中心,清华大学)

Comments Accepted by ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04410 2025-07-08 cs.CV cs.AI cs.IR

Multimedia Verification Through Multi-Agent Deep Research Multimodal Large Language Models

Huy Hoan Le, Van Sy Thinh Nguyen, Thi Le Chi Dang, Vo Thanh Khang Nguyen, Truong Thanh Hung Nguyen, Hung Cao

机构 * University of New Brunswick(新 Brunswick大学) Institute for Clarity in Documentation(文档清晰研究所) Inria Paris-Rocquencourt(巴黎-罗克琴特研究所) Rajiv Gandhi University(拉贾·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒尔研究实验室)

Comments 33rd ACM International Conference on Multimedia (MM'25) Grand Challenge on Multimedia Verification

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03589 2025-07-08 cs.CV cs.AI cs.CL

BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance

Huy Le, Nhat Chung, Tung Kieu, Anh Nguyen, Ngan Le

机构 * FPT Software AI Center(FPT软件AI中心) Aalborg University(奥尔堡大学) Pioneer Centre for AI(先锋人工智能中心) University of Liverpool(利物浦大学) AICV Lab, University of Arkansas(AICV实验室,阿肯色大学)

Comments Accepted at ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20948 2025-07-08 cs.CV

DS_FusionNet: Dynamic Dual-Stream Fusion with Bidirectional Knowledge Distillation for Plant Disease Recognition

Yanghui Song, Chengfu Yang

Comments 9 pages, 14 figures, 10th International Conference on Computer-Aided Design, Manufacturing, Modeling and Simulation (CDMMS 2025)

Journal ref Proceedings of the 32nd ACM International Conference on Multimedia. 2024: 1593-1601

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00606 2025-07-08 cs.CV

Sample-level Adaptive Knowledge Distillation for Action Recognition

Ping Li, Chenhao Ping, Wenxiao Wang, Mingli Song

Journal ref ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.01819 2025-07-08 cs.CV cs.AI

Free-Mask: A Novel Paradigm of Integration Between the Segmentation Diffusion Model and Image Editing

Bo Gao, Jianhui Wang, Xinyuan Song, Yangfan He, Fangxu Xing, Tianyu Shi

机构 * Sun Yat-sen University(中山大学) University of Electronic Science and Technology of China(电子科学与技术大学) Emory University(埃默里大学) University of Minnesota-Twin Cities(明尼苏达大学双城分校) Harvard Medical School(哈佛医学院) University of Toronto(多伦多大学)

Comments Accepted by ACM MM(2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10428 2025-07-01 cs.CV

RCA: Region Conditioned Adaptation for Visual Abductive Reasoning

Hao Zhang, Yeo Keat Ee, Basura Fernando

机构 * Centre for Frontier AI Research, Agency for Science, Technology and Research (A*STAR)(前沿人工智能研究中心,科技研究局(A*STAR))

Comments 13 pages, 11 figures, ACM Multimedia 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.02157 2025-06-25 cs.CV cs.HC

FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs

Haodong Chen, Haojian Huang, Junhao Dong, Mingzhe Zheng, Dian Shao

机构 * School of Automation, Northwestern Polytechnical University(西北工业大学自动化学院) The University of Hong Kong(香港大学) Nanyang Technological University(南洋理工大学) School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机学院) Unmanned System Research Institute, Northwestern Polytechnical University(西北工业大学无人系统研究所)

Comments Accepted to ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10443 2025-06-13 cs.LG

MNN-LLM: A Generic Inference Engine for Fast Large Language Model Deployment on Mobile Devices

Zhaode Wang, Jingbang Yang, Xinyu Qian, Shiwen Xing, Xiaotang Jiang, Chengfei Lv, Shengyu Zhang

机构 * Alibaba Group(阿里巴巴集团) Zhejiang University(浙江大学)

Comments 7 pages, 5 figures. Published in the Proceedings of the 6th ACM International Conference on Multimedia in Asia Workshops (MMAsia '24 Workshops). The final authenticated version is available at https://dl.acm.org/doi/10.1145/3700410.3702126

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17152 2025-06-12 cs.CV cs.AI

XMeCap: Meme Caption Generation with Sub-Image Adaptability

Yuyan Chen, Songzhou Yan, Zhihong Zhu, Zhixu Li, Yanghua Xiao

机构 * Shanghai Key Laboratory of Data Science, School of Computer Science, Fudan University(复旦大学计算机学院数据科学实验室) Peking University(北京大学)

Comments Accepted to ACM Multimedia 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11082 2025-06-10 cs.CV cs.AI cs.CL cs.MM

Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive Learning

Chen Jiang, Hong Liu, Xuzheng Yu, Qing Wang, Yuan Cheng, Jia Xu, Zhongyi Liu, Qingpei Guo, Wei Chu, Ming Yang, Yuan Qi

机构 * Artificial Intelligence Innovation and Incubation Institute, Fudan University(复旦大学人工智能创新与孵化院) Ant Group(蚂蚁集团)

Comments Accepted by ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12274 2025-06-05 cs.CV

MDPE: A Multimodal Deception Dataset with Personality and Emotional Characteristics

Cong Cai, Shan Liang, Xuefei Liu, Kang Zhu, Zhengqi Wen, Jianhua Tao, Heng Xie, Jizhou Cui, Yiming Ma, Zhenhua Cheng, Hanzhe Xu, Ruibo Fu, Bin Liu, Yongwei Li

机构 * Beijing Institute of Technology(北京理工大学) Xi’an Jiaotong Liverpool University(西安交通大学利物浦大学) Institute of Automation, Chinese Academy of Sciences (CAS)(中国科学院自动化研究所) Anhui University(安徽大学) Beijing National Research Center for Information Science and Technology, Tsinghua University(北京信息科学与技术国家研究中心,清华大学) Department of Automation, Tsinghua University(清华大学自动化系) ShanghaiTech University(上海科技大学) University of Chinese Academy of Sciences(中国科学院大学) Tianjin Normal University(天津师范大学) Institute of Automation, CAS(中国科学院自动化研究所) Institute of Psychology, CAS(中国科学院心理研究所)

Comments Code and data are available; Submitted to ACM Multimedia 2025 Dataset Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10034 2025-05-30 cs.AI

The First MPDD Challenge: Multimodal Personality-aware Depression Detection

Changzeng Fu, Zelin Fu, Qi Zhang, Xinhe Kuang, Jiacheng Dong, Kaifeng Su, Yikai Su, Wenbo Shi, Junfeng Yao, Yuliang Zhao, Shiqi Zhao, Jiadong Wang, Siyang Song, Chaoran Liu, Yuichiro Yoshikawa, Björn Schuller, Hiroshi Ishiguro

机构 * Northeastern University(东北大学) University of Technology Sydney(悉尼大学) Xiamen University(厦门大学) Technical University of Munich(慕尼黑技术大学) University of Cambridge(剑桥大学) National Information Institute(国家信息研究所) Osaka University(大阪大学) Imperial College London(伦敦帝国理工学院)

Comments This paper has been accepted as part of the MPDD Challenge in the ACMMM 2025 Grand Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.08190 2025-05-27 cs.CV cs.LG

Graph Neural Networks for Knowledge Enhanced Visual Representation of Paintings

Athanasios Efthymiou, Stevan Rudinac, Monika Kackovic, Marcel Worring, Nachoem Wijnberg

机构 * University of Amsterdam(阿姆斯特丹大学)

Comments Published in the 29th ACM International Conference on Multimedia (MM '21). This is the camera-ready version. 10 pages, 4 figures

Journal ref Proc. 29th ACM Int. Conf. on Multimedia (MM '21), 2021, pp. 3710-3719

详情

展开后加载摘要…

URL PDF HTML 收藏