arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

ACM International Conference on Multimedia · 会议 · Multimedia

2026-08-28 至 2026-08-28 共收录 6
2608.27214 2026-08-28 cs.CV 新提交

CODE: Cross-Modal Calibration and Dynamic Suppression for Open World Object Detection

CODE:面向开放世界目标检测的跨模态校准与动态抑制

Hao Xu, Zhaoning Shi, Hehe Jin, Bo Ma

机构 * Beijing Institute of Technology(北京理工大学)

AI总结 本文针对开放世界目标检测的语义歧义与过度抑制问题,提出含三个互补组件的CODE框架,在真实世界检测基准上超越此前最优方法。

Comments Accepted by ACM Multimedia 2026 (MM '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.27127 2026-08-28 cs.AI 新提交

TransMeme: A Multi-Agent Framework for Cross-Cultural Meme Transcreation

TransMeme:用于跨文化梗图再创作的多智能体框架

Jingyi Zheng, Yule Liu, Zifan Peng, Tianyi Hu, Yuemeng Zhao, Xinhu Zheng, Xinlei He

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Wuhan University(武汉大学) Aarhus University(奥胡斯大学)

AI总结 本研究针对跨文化梗图再创作的三大核心挑战,提出多智能体框架TransMeme,经中英双向梗图再创作的人工评估与LLM评判,性能优于所有基线,为该领域未来幽默迁移研究指明方向。

Comments 10 pages, 4 figures. Accepted at the 34th ACM International Conference on Multimedia (ACM MM 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26971 2026-08-28 cs.CV cs.MM 新提交

TempJail: Temporal Jailbreak Attacks against Image-to-Video Generation Models

TempJail:针对图像到视频生成模型的时序越狱攻击

Qi Lu, Zehui Guo, David Yuanda Gan, Zijing Li, Hengda Zhang, Weijun Xu, Qiankun Zhang

机构 * School of Cyber Science and Engineering, Huazhong University of Science and Technology(华中科技大学网络空间科学与工程学院) School of Mathematical Sciences, Peking University(北京大学数学科学学院) School of Software and engineering, Huazhong University of Science and Technology(华中科技大学软件学院) School of Computer Science, Nanjing University(南京大学计算机学院)

AI总结 本文提出TempJail时序越狱框架,针对图像到视频生成模型的时序漏洞,通过分解恶意提示、受控潜在扰动等方式提升攻击成功率,在Kling等模型上获显著效果。

Comments Accepted by ACM Multimedia 2026 (ACM MM '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26820 2026-08-28 cs.CV 新提交

LLaVAFlow: Preserving Latent Alignment Flow for Parameter-Efficient Multimodal Fine-Tuning

LLaVAFlow:用于参数高效多模态微调的潜在对齐流保留方法

Muyao Yuan, Muyan Jiao, Jiangyong Ying, Weizhan Zhang, Yuanhong Zhang, Lan Ma, Yuan Gao, Haipeng Du

机构 * MOEKLINNS China Telecom(中国电信)

AI总结 针对多模态大语言模型微调的灾难性遗忘问题,本文提出即插即用的LLaVAFlow框架,通过信息论蒸馏保留跨模态对齐流,提升下游任务性能与泛化能力。

Comments Accepted by ACM Multimedia 2026 (ACM MM 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26752 2026-08-28 cs.CV 新提交

Glass Surface Detection Grounded in 3D Visual Geometry

基于3D视觉几何的玻璃表面检测

Yiwei Lu, Ke Xu, Tao Yan, Xiaojun Chang, Radu Timofte, Rynson W. H. Lau

机构 * Jiangnan University(江南大学) University of Science and Technology of China(中国科学技术大学) University of Wurzburg(维尔茨堡大学) City University of Hong Kong (Dongguan)(香港城市大学(东莞))

AI总结 本文提出将玻璃表面检测(GSD)基于3D视觉几何的新范式,结合VGGT、FSAM与GeGB,在7个基准上达最优性能,泛化性好且提升玻璃场景重建效果。

Comments 9 pages, 10 figures. Accepted by ACM Multimedia 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26744 2026-08-28 cs.CV 新提交

G2D: Generative-to-Discriminative Collaborative Inference for Zero-Shot Image Classification

G2D:用于零样本图像分类的生成式到判别式协同推理框架

Zehua Hao, Fang Liu, Qinliang Wang, Yaoyang Du, Xinyan Huang, Puhua Chen

机构 * Xidian University(西安电子科技大学)

AI总结 G2D是一种无需训练的零样本图像分类框架,通过生成式VLM验证CLIP检索的候选,在8个基准上平均准确率达68.85%,优于CLIP及独立生成式模型,还可迁移至其他模型。

Comments Accepted at ACM MM 2026. 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏