arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

ACM International Conference on Multimedia · 会议 · Multimedia

2026-08-06 至 2026-08-06 共收录 8
2608.04891 2026-08-06 eess.SP 新提交

GVC-RT: Towards Real-Time Generative Video Compression at Ultra-Low Bitrates

GVC-RT:面向超低码率下的实时生成视频压缩

Tianjian Dang, Sixian Wang, Lei Luo, Guo Lu, Jincheng Dai

AI总结 本文针对现有生成视频编解码器实时性不足的问题,提出GVC-RT,通过重新设计生成式隐编码框架,在超低码率下实现了优于SOTA模型的压缩性能与实时编码解码速度。

Comments Accepted to appear in the Proceedings of the 34th ACM International Conference on Multimedia (MM '26). 10 pages, 9 figures, and 2 tables. Code: https://github.com/semcomm/GVC-RT

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05126 2026-08-06 cs.CL cs.MM 新提交

Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models

语音函数调用:面向大型音频语言模型的语音理解新视角

Yuezhang Peng, Yuxin Liu, Changfeng Gao, Zhifu Gao, Xiangang Li, Xie Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Token Foundry, Alibaba Group(阿里巴巴集团Token Foundry) Shanghai Innovation Institute(上海创新研究院)

AI总结 该研究提出语音函数调用(SFC)这一新型语义理解视角,构建SFC-Bench数据集并评估LLMs与LALMs性能,经后训练提升LALMs的SFC能力,实验显示SFC优于传统SLU,可大幅提升语义提取准确率。

Comments ACM Multimedia 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05101 2026-08-06 cs.CV cs.MM 新提交

HexMIL: Hierarchical Attention MIL for Ante-Hoc Explainable Detection of AI-Manipulated CT Volumes

HexMIL:用于AI篡改CT体积事前可解释检测的分层注意力多实例学习

Orazio Pontorno, Luca Guarnera, Zahid Akhtar, Sebastiano Battiato

机构 * University of Catania(卡塔尼亚大学) State University of New York Polytechnic Institute(纽约州立大学理工学院)

AI总结 HexMIL是一种无掩码的医疗深度伪造检测器,采用分层注意力多实例学习,利用二元体积级监督实现AI篡改CT体积的事前可解释检测,在跨生成器泛化任务中性能优于基线。

Comments Accepted at ACM Multimedia 2026 (MM '26)

Journal ref Proceedings of the 34th ACM International Conference on Multimedia (MM '26), November 10--14, 2026, Rio de Janeiro, Brazil

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04750 2026-08-06 cs.CV cs.CL cs.MM 新提交

Simile Understanding in Text-to-Image Models: An Evaluation Framework

文本到图像模型中的明喻理解:一个评估框架

Luecheng Wang, Shintaro Ozaki, Hidetaka Kamigaito, Katsuhiko Hayashi, Jingun Kwon, Manabu Okumura, Taro Watanabe

机构 * The University of Tokyo(东京大学) Nara Institute of Science and Technology(奈良科学技术研究所) Chungnam National University(忠南国立大学) Institute of Science Tokyo(东京科学大学)

AI总结 针对文本到图像模型常混淆明喻喻体与本体的问题,提出含受控数据集、YOLO指标及Diffusion Lens分析的评估框架,实验发现模型存在字面化失败模式并讨论了缓解策略。

Comments Accepted as a full paper at ACM Multimedia 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04655 2026-08-06 cs.CV cs.AI 新提交

CSGen: A Multi-Domain Curvilinear Structure Generation Model via Hierarchical Multimodal Diffusion

CSGen:一种基于分层多模态扩散的多领域曲线结构生成模型

Zhe Shan, Ziming Yang, Lei Zhou, Wenwen Zhang, Cong Lin, Xia Xie

机构 * Hainan University(海南大学) Guangdong Ocean University(广东海洋大学)

AI总结 该研究针对可控曲线结构图像生成的挑战,提出CSGen模型,通过构建多领域数据集、分层渐进控制策略与稀疏感知损失机制,提升生成图像的结构精度与下游分割性能,为曲线结构分析提供新范式。

Comments Accepted to ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04610 2026-08-06 cs.CV 新提交

HiSC: Hierarchical Spatial Clustering Token Compression for Efficient 3D Scene Understanding

HiSC:用于高效3D场景理解的分层空间聚类令牌压缩

Jiuhe Qu, Yingping Liang, Ying Fu

机构 * Beijing Institute of Technology(北京理工大学)

AI总结 本文提出无训练框架HiSC,通过SGraM策略和SCluP范式实现3D VLMs的分层空间聚类令牌压缩,在高剪枝率下实现超90%令牌减少且性能下降极小,验证了其有效性。

Comments Accepted by ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04581 2026-08-06 cs.CV 新提交

ACA-GS: Adaptive-Capacity Anchored Gaussian Splatting for Compact Dynamic Radiance Fields

ACA-GS:用于紧凑动态辐射场的自适应容量锚定高斯溅射

Seunghyeon Song, Joo Chan Lee, Chanung Park, Jun Young Jeong, Minseo Lee, Eunbyung Park, Jong Hwan Ko

机构 * Sungkyunkwan University(成均馆大学) Electronics and Telecommunications Research Institute(电子通信研究院) Yonsei University(延世大学)

AI总结 针对4D高斯溅射中运动表达与存储效率的权衡,提出自适应容量锚定高斯溅射框架,通过调整锚点的神经高斯数量和特征通道实现紧凑动态辐射场,在多数据集上显著压缩存储且不降低质量。

Comments 9 pages, 8 figures. Accepted to ACM Multimedia 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04385 2026-08-06 cs.CV 新提交

ReGround: Restoring Visual Grounding in Multi-Step Reasoning through Self-Diagnosis and Visual Re-Examination

ReGround:通过自诊断与视觉重检验恢复多步推理中的视觉接地

Lei Peng, Shuai Lv, Wei Hu

机构 * University of Science and Technology of China(中国科学技术大学) School of Artificial Intelligence and Data Science(人工智能与数据科学学院) State Key Laboratory of Precision and Intelligent Chemistry(精准与智能化学国家重点实验室)

AI总结 ReGround是无需架构修改或外部工具的两阶段框架,通过自诊断与视觉重检验解决VLMs多步推理中的视觉接地丢失问题,在八个基准上获一致增益且推理开销适度。

Comments Accepted to ACM Multimedia 2026 (MM '26). 8 pages main text, 4 figures, plus appendix

详情

展开后加载摘要…

URL PDF HTML 收藏