arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Huazhong University of Science and Technology(华中科技大学)

2026-09-01 至 2026-09-01 共收录 6
2608.29958 2026-09-01 cs.CV 新提交

RIDGE: Region-Informed Derivative-Guided Evidence Selection for Long Video Understanding

RIDGE:用于长视频理解的区域感知导数引导证据选择

Shanqing Xu, Meng Luo, Mengchen Qian, Yuhui Gao, Siyue Peng, Xiaohan Zhong, Xiaojin Zhang, Zhongyu Wei, Wei Chen, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) National University of Singapore(新加坡国立大学) Fudan University(复旦大学)

AI总结 针对长视频理解中视觉内容超出LVLM固定token预算的问题,提出RIDGE框架,通过将帧-查询相似性曲线作为时间信号划分区域并选择证据,在多基准和主干上取得最优性能。

Comments EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29865 2026-09-01 cs.CL 新提交

ManGo: Manga Active Narrative Grounding Optimization

ManGo:漫画主动叙事定位优化

Hao Qiu, Junyan Wang, Zheyuan Liu, Lei Fan, Hong Jia, Lianbo Guo, Zhulin Tao

机构 * Communication University of China(中国传媒大学) Australian Institute for Machine Learning, Adelaide University(阿德莱德大学澳大利亚机器学习研究所) University of New South Wales(新南威尔士大学) University of Auckland(奥克兰大学) Huazhong University of Science and Technology(华中科技大学)

AI总结 本研究针对漫画视觉问答的分镜叙事结构问题,提出无监督框架ManGo,通过主动叙事草图结合双奖励的组相对策略训练,在标准基准上取得最优性能。

Comments 16 pages, 9 figures, 7 tables. Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29233 2026-09-01 cs.CV 新提交

Generalization over Memorization: Generalization-Aware Diffusion Adaptation for Single-Image Multi-View Synthesis

超越记忆的泛化:面向单图像多视图合成的泛化感知扩散适配

Jie Li, Xingchen Zou, Yuxuan Liang

机构 * Huazhong University of Science and Technology(华中科技大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

AI总结 本研究提出ACM MM 2026单图像多视图合成挑战赛的获奖方案GoM,通过场景不相交验证等技术,在40个训练场景下实现293支队伍中排名第一,揭示小数据生成建模中验证设计与训练轨迹控制的关键作用。

Comments ACM Multimedia 2026 Grand Challenge Track winning paper; presents the 1st-place solution among 293 registered teams. 7 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29160 2026-09-01 cs.CV cs.AI 新提交

Training-Free Hidden-State Refinement for Flow-Matching Image Generators

无需训练的流匹配图像生成器的隐状态优化

Yuanyi Yan, Xinzhe Rao, Canyu Shen, Yang Chen, Yunlu Chen, Meng Tang, Teng Long, Vincent Tao Hu

机构 * Huazhong University of Science and Technology(华中科技大学) Tongji University(同济大学) King Abdullah University of Science and Technology(阿卜杜拉国王科技大学) University of California, Merced(加州大学默塞德分校) University of Amsterdam(阿姆斯特丹大学)

AI总结 该研究提出无需训练的循环框架优化流匹配图像生成器,在不改动权重与采样器的情况下提升质量指标,Loop Guidance在Scale-RAE DiT2.4B上显著提升GenEval与DPG-Bench指标。

Comments 7pages,4 figures,5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29113 2026-09-01 cs.CV 新提交

GramLoop: Training-Free Gram-Gated Replay for Robust Dense Prediction

GramLoop:用于鲁棒密集预测的无训练Gram门控重放方法

Yang Chen, Canyu Shen, Xinzhe Rao, Yuanyi Yan, Yunlu Chen, Meng Tang, Teng Long, Vincent Tao Hu

机构 * Huazhong University of Science and Technology(华中科技大学) Tongji University(同济大学) King Abdullah University of Science and Technology(阿卜杜拉国王科技大学) University of California, Merced(加州大学默塞德分校) University of Amsterdam(阿姆斯特丹大学)

AI总结 该研究提出无训练框架GramLoop,通过在冻结DINOv3的视觉骨干内添加推理计算,在不改变模型结构的情况下,提升了分布偏移下目标检测和语义分割等密集预测任务的鲁棒性,在COCO-O等基准上取得性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.28687 2026-09-01 cs.CV 新提交

FLM: Frequency-Aware Language Models for Generative Image Compression

FLM:用于生成式图像压缩的频率感知语言模型

Jiarun Chen, Kejun Wu, Li Li, Chengtao Cai, Zhengguo Li, Chia-Wen Lin

机构 * School of Electronic Information and Communications, Huazhong University of Science and Technology(华中科技大学电子信息与通信学院) University of Science and Technology of China(中国科学技术大学) College of Intelligent Systems Science and Engineering, Harbin Engineering University(哈尔滨工程大学智能系统科学与工程学院) Institute for Infocomm Research, Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局信息通信研究院) National Tsing Hua University(国立清华大学)

AI总结 本研究提出频率感知语言模型FLM,通过频域概率建模实现高效图像压缩,兼容有损与无损JPEG重压缩,在多数据集上较JPEG基准获显著BD-PSNR增益,性能优于传统及生成式压缩方法。

详情

展开后加载摘要…

URL PDF HTML 收藏