arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

ACM International Conference on Multimedia · 会议 · Multimedia

共收录 1944
2507.20326 2025-07-29 cs.LG cs.AI

MIPS: a Multimodal Infinite Polymer Sequence Pre-training Framework for Polymer Property Prediction

Jiaxi Wang, Yaosen Min, Xun Zhu, Miao Li, Ji Wu

机构 * Department of Electronic Engineering, Tsinghua University Beijing China Zhongguancun Institute of Artificial Intelligence Beijing China Department of Electronic Engineering \& College of AI, Tsinghua University Beijing National Research Center for Information Science Department of Electronic Engineering, Tsinghua University Zhongguancun Institute of Artificial Intelligence

Comments 14 pages, 8 figures, accepted by ACM Multimedia 2025 (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19949 2025-07-29 cs.CV

AF-CLIP: Zero-Shot Anomaly Detection via Anomaly-Focused CLIP Adaptation

Qingqing Fang, Wenxi Lv, Qinliang Su

机构 * School of Computer Science and Engineering(计算机科学与工程学院) Sun Yat-sen University(中山大学) Guangdong Key Laboratory of Big Data Analysis and Processing(大数据分析与处理广东省重点实验室)

Comments The paper is accepted by ACM MM' 25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19863 2025-07-29 cs.MM

Anchoring Trends: Mitigating Social Media Popularity Prediction Drift via Feature Clustering and Expansion

Chia-Ming Lee, Bo-Cheng Qiu, Cheng-Jun Kang, Yi-Hsuan Wu, Jun-Lin Chen, Yu-Fan Lin, Yi-Shiuan Chou, Chih-Chung Hsu

Comments Accepted by ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19836 2025-07-29 cs.GR cs.AI cs.CV cs.MM cs.SD

ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion

Xuanchen Wang, Heng Wang, Weidong Cai

机构 * School of Computer Science, The University of Sydney(计算机科学学院,悉尼大学)

Comments 10 pages, 5 figures, accepted by the 33rd ACM International Conference on Multimedia (ACM MM 2025), demo page: https://choreomuse.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19835 2025-07-29 cs.SD cs.MM

SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations

Chunshi Wang, Hongxing Li, Yawei Luo

机构 * Zhejiang University(浙江大学)

Comments Accepted by ACMMM'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19807 2025-07-29 cs.CV

DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection

Guiping Cao, Xiangyuan Lan, Wenjian Huang, Jianguo Zhang, Dongmei Jiang, Yaowei Wang

机构 * RITAS, Southern University of Science and Technology(南方科技大学RITAS) Pengcheng Laboratory(鹏城实验室) Pazhou Laboratory (Huangpu)(黄埔实验室) Department of Computer Science and Engineering, Southern University of Science and Technology(南方科技大学计算机科学与工程系) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

Journal ref The 33rd ACM International Conference on Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15765 2025-07-29 cs.CV cs.AI

Learning from Heterogeneity: Generalizing Dynamic Facial Expression Recognition via Distributionally Robust Optimization

Feng-Qi Cui, Anyang Tong, Jinyang Huang, Jie Zhang, Dan Guo, Zhi Liu, Meng Wang

机构 * Hefei University of Technology(合肥工业大学) The University of Electro-Communications(电通大学)

Comments Accepted by ACM MM'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13092 2025-07-29 cs.CV

EventVAD: Training-Free Event-Aware Video Anomaly Detection

Yihua Shao, Haojin He, Sijie Li, Siyu Chen, Xinwei Long, Fanhu Zeng, Yuxuan Fan, Muyang Zhang, Ziyang Yan, Ao Ma, Xiaochen Wang, Hao Tang, Yan Wang, Shuyan Li

机构 * Peking University(北京大学) Guangdong University of Technology(广东工业大学) The University of Sheffield(谢菲尔德大学) University of Science and Technology Beijing(北京科技大学) Tsinghua University(清华大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Nanjing University(南京大学) University of Trento(特伦特大学) Queen's University Belfast(贝尔法斯特女王大学)

Comments Paper was accepted by ACM MM 2025; Code: https://github.com/YihuaJerry/EventVAD

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15278 2025-07-29 cs.CV cs.AI

CopyJudge: Automated Copyright Infringement Identification and Mitigation in Text-to-Image Diffusion Models

Shunchang Liu, Zhuan Shi, Lingjuan Lyu, Yaochu Jin, Boi Faltings

机构 * EPFL(苏黎世联邦理工学院) Mila - Quebec AI Institute Mcgill University(蒙特利尔麦吉尔大学人工智能研究所) Westlake University(西湖大学)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19201 2025-07-28 eess.IV cs.AI cs.CV

Joint Holistic and Lesion Controllable Mammogram Synthesis via Gated Conditional Diffusion Model

Xin Li, Kaixiang Yang, Qiang Li, Zhiwei Wang

机构 * Wuhan National Laboratory for Optoelectronics, Huazhong University of Science and Technology(华中科技大学光电研究院)

Comments Accepted, ACM Multimedia 2025, 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19077 2025-07-28 cs.CV

Multi-Task Dense Prediction Fine-Tuning with Mixture of Fine-Grained Experts

Yangyang Xu, Xi Ye, Duo Su

机构 * Department of Computer Science and Technology, Tsinghua University(计算机科学与技术系,清华大学)

Comments Accepted to ACM Multimedia 2025 (MM'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19062 2025-07-28 cs.SD eess.AS

From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language Models

Zhaoxi Mu, Rilin Chen, Andong Li, Meng Yu, Xinyu Yang, Dong Yu

机构 * Xi'an Jiaotong University(西安交通大学) Tencent AI Lab(腾讯AI实验室) Institute of Acoustics, Chinese Academy of Sciences(中国科学院声学研究所)

Comments ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18881 2025-07-28 cs.CV cs.RO

Perspective from a Higher Dimension: Can 3D Geometric Priors Help Visual Floorplan Localization?

Bolei Chen, Jiaxu Kang, Haonan Yang, Ping Zhong, Jianxin Wang

机构 * School of Computer Science Engineering, Central South University Changsha Hunan China Engineering, Central South University

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09540 2025-07-28 cs.CV

EmbodiedOcc++: Boosting Embodied 3D Occupancy Prediction with Plane Regularization and Uncertainty Sampler

Hao Wang, Xiaobao Wei, Xiaoan Zhang, Jianing Li, Chengyu Bai, Ying Li, Ming Lu, Wenzhao Zheng, Shanghang Zhang

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学) School of electronic science and engineering, Nanjing University(电子科学与工程学院,南京大学) Berkeley Artificial Intelligence Research Lab, Department of EECS, University of California, Berkeley(伯克利人工智能研究实验室,电子工程与计算机科学系,加州大学伯克利分校)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15775 2025-07-28 cs.CV cs.SE

Do Existing Testing Tools Really Uncover Gender Bias in Text-to-Image Models?

Yunbo Lyu, Zhou Yang, Yuqing Niu, Jing Jiang, David Lo

机构 * Singapore Management University(新加坡管理大学) University of Alberta(阿尔伯塔大学) Australian National University(澳大利亚国立大学)

Comments Accepted to ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18616 2025-07-25 cs.CV cs.AI cs.CL cs.LG

SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning

Si-Woo Kim, MinJu Jeon, Ye-Chan Kim, Soeun Lee, Taewhan Kim, Dong-Jin Kim

机构 * Hanyang University(翰阳大学)

Comments Accepted to ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18243 2025-07-25 cs.CV cs.AI

DepthDark: Robust Monocular Depth Estimation for Low-Light Environments

Longjian Zeng, Zunjie Zhu, Rongfeng Lu, Ming Lu, Bolun Zheng, Chenggang Yan, Anke Xue

机构 * Hangzhou Dianzi University(杭州电子科技大学) Intel Labs China(英特尔中国实验室)

Comments Accepted by ACM MM 2025 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04600 2025-07-25 cs.AI

DisMS-TS: Eliminating Redundant Multi-Scale Features for Time Series Classification

Zhipeng Liu, Peibo Duan, Binwu Wang, Xuan Tang, Qi Chu, Changsheng Zhang, Yongsheng Huang, Bin Zhang

机构 * School of Software, Northeastern University(东北大学软件学院) School of Software, University of Science and Technology of China(中国科学技术大学软件学院) School of Software, University of Science(科学大学软件学院)

Comments This paper has been accepted for presentation at the ACM International Conference on Multimedia (ACM MM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.00970 2025-07-25 cs.MM

Multimodal Fusion via Hypergraph Autoencoder and Contrastive Learning for Emotion Recognition in Conversation

Zijian Yi, Ziming Zhao, Zhishu Shen, Tiehua Zhang

Comments Accepted by ACM MULTIMEDIA 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17687 2025-07-24 cs.LG

Towards Effective Open-set Graph Class-incremental Learning

Jiazhen Chen, Zheng Ma, Sichao Fu, Mingbin Feng, Tony S. Wirjanto, Weihua Ou

机构 * University of Waterloo(滑铁卢大学) Huazhong University of Science and Technology(华中科技大学) Guizhou Normal University(贵州师范大学)

Comments Accepted by 33rd ACM International Conference on Multimedia (MM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17456 2025-07-24 cs.CV

Dynamic Scoring with Enhanced Semantics for Training-Free Human-Object Interaction Detection

Francesco Tonini, Lorenzo Vaquero, Alessandro Conti, Cigdem Beyan, Elisa Ricci

机构 * University of Trento(特伦托大学) University of Verona(威尼斯大学) Department of Computer Science(计算机科学系)

Comments Accepted to ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09184 2025-07-24 cs.CV

MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models

Qiyan Zhao, Xiaofeng Zhang, Yiheng Li, Yun Xing, Xiaosong Yuan, Feilong Tang, Sinan Fan, Xuhang Chen, Xuyao Zhang, Dahan Wang

机构 * FKLPRIU, Xiamen University of Technology, China(福克斯理工国际大学,厦门理工学院,中国) Shanghai Jiao Tong University, China(上海交通大学,中国) Nanyang Technological University, Singapore(南洋理工大学,新加坡) Jilin University, China(吉林大学,中国) Monash University, Australia(墨尔本大学,澳大利亚) Zhejiang University, China(浙江大学,中国) Huizhou University, China(惠州大学,中国) Chinese Academy of Sciences, China(中国科学院,中国)

Comments Accepted in ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07578 2025-07-24 cs.CV

Diffusion-Guided Knowledge Distillation for Weakly-Supervised Low-Light Semantic Segmentation

Chunyan Wang, Dong Zhang, Jinhui Tang

机构 * Nanjing University of Science and Technology(南京理工大学) The Hong Kong University of Science and Technology(香港科学与技术大学) Nanjing Forestry University(南京林业大学)

Comments Accepted by ACM Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16619 2025-07-24 cs.CV

Human-Activity AGV Quality Assessment: A Benchmark Dataset and an Objective Evaluation Metric

Zhichao Zhang, Wei Sun, Xinyue Li, Yunhao Li, Qihang Ge, Jun Jia, Zicheng Zhang, Zhongpeng Ji, Fengyu Sun, Shangling Jui, Xiongkuo Min, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) Huawei Technologies(华为技术)

Comments Accepted by ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16472 2025-07-23 cs.CV

DenseSR: Image Shadow Removal as Dense Prediction

Yu-Fan Lin, Chia-Ming Lee, Chih-Chung Hsu

机构 * National Cheng Kung University(国立成功大学) National Yang Ming Chiao Tung University(国立阳明交通大学)

Comments Paper accepted to ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16238 2025-07-23 cs.CV

Positive Style Accumulation: A Style Screening and Continuous Utilization Framework for Federated DG-ReID

Xin Xu, Chaoyue Ren, Wei Liu, Wenke Huang, Bin Yang, Zhixi Yu, Kui Jiang

机构 * Wuhan University of Science(武汉科技大学) Wuhan University(武汉大学) Harbin Institute of Technology(哈尔滨工业大学)

Comments 10 pages, 3 figures, accepted at ACM MM 2025, Submission ID: 4394

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15253 2025-07-22 cs.AI cs.LG cs.SI

Disentangling Homophily and Heterophily in Multimodal Graph Clustering

Zhaochen Guo, Zhixiang Shen, Xuanting Xie, Liangjian Wen, Zhao Kang

机构 * University of Electronic Science and Technology of China(电子科技大学) Southwestern University of Finance and Economics(西南财经大学)

Comments Appear in ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20090 2025-07-22 cs.CV

Transfer Attack for Bad and Good: Explain and Boost Adversarial Transferability across Multimodal Large Language Models

Hao Cheng, Erjia Xiao, Jiayan Yang, Jinhao Duan, Yichi Wang, Jiahang Cao, Qiang Zhang, Le Yang, Kaidi Xu, Jindong Gu, Renjing Xu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of Oxford(牛津大学) Drexel University(德雷塞尔大学) Beijing University of Technology(北京理工大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Xi’an Jiaotong University(西安交通大学)

Comments This paper is accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10888 2025-07-21 cs.CV cs.AI

CDUPatch: Color-Driven Universal Adversarial Patch Attack for Dual-Modal Visible-Infrared Detectors

Jiahuan Long, Wen Yao, Tingsong Jiang, Chao Ma

机构 * Chinese Academy of Military Science(中国军事科学院) Shanghai Jiao Tong University(上海交通大学)

Comments Accepted by ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19243 2025-07-21 cs.CV

Accelerating Diffusion Transformer via Error-Optimized Cache

Junxiang Qiu, Shuo Wang, Jinda Lu, Lin Liu, Houcheng Jiang, Xingyu Zhu, Yanbin Hao

机构 * University of Science and Technology of China(中国科学技术大学) Hefei University of Technology(合肥工业大学)

Journal ref ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏