arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

非对称跨模态细粒度视觉分类:ACF-Net与BirdPro基准

Asymmetric Cross-Modal Fine-Grained Visual Categorization: ACF-Net and the BirdPro Benchmark

Bohan Deng, Shuo Ye, Zitong Yu

arXiv 2608.25520首次发表:更新:

发表机构

Great Bay University(大湾区大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对非对称跨模态细粒度视觉分类的挑战,提出ACF-Net框架并构建BirdPro基准,实验显示其在融合与不匹配设置中均优于基线方法。

AI 中文摘要

视听跨模态细粒度视觉分类(FGVC)旨在通过联合利用视觉和听觉信息识别细粒度类别。然而,非对称跨模态场景下的FGVC受到的关注有限,该场景中配对的视频和音频并非严格同步,甚至可能不对应同一主体或时刻。这种薄弱且模糊的跨模态对应关系对有效表示学习和模态对齐构成了重大挑战。为解决这些问题,我们提出ACF-Net,一种新型光流引导的非对称视听细粒度学习框架。ACF-Net包含两个关键模块:光流引导运动(OFGM)和非对称跨模态自适应融合(ACAF)。OFGM捕捉对运动敏感的视觉线索并抑制无关背景干扰,从而增强视频中具有判别性的动态表示。ACAF在弱匹配的视听对下估计模态可靠性,并执行感知不确定性的自适应融合,以提升类别级识别的鲁棒性。为支持非对称跨模态FGVC的研究,我们进一步构建了BirdPro,一个面向鸟类的新型视听基准,因为现有数据集往往缺乏非严格时间和实例对应关系下的大规模类别级视听关联。BirdPro包含1919条音频记录和11965个视频,覆盖194种鸟类。大量实验表明,与代表性基线方法相比,ACF-Net取得了最佳结果,在融合设置和不匹配设置中分别比最强基线高出2.97%和1.92%。

英文摘要

Audio-visual cross-modal Fine-Grained Visual Categorization (FGVC) aims to identify fine-grained categories by jointly leveraging visual and auditory information. However, FGVC under asymmetric cross-modal scenarios has received limited attention, where paired video and audio are not strictly synchronized and may not even correspond to the same individual or moment. Such weak and ambiguous cross-modal correspondence poses substantial challenges to effective representation learning and modality alignment. To address these issues, we propose ACF-Net, a novel optical flow-guided framework for asymmetric audio-visual fine-grained learning. ACF-Net consists of two key modules: Optical Flow-Guided Motion (OFGM) and Asymmetric CrossModal Adaptive Fusion (ACAF). OFGM captures motion-sensitive visual cues and suppresses irrelevant background interference, thereby enhancing discriminative dynamic representations in videos. ACAF estimates modality reliability under weakly matched audio-video pairs and performs uncertainty-aware adaptive fusion to improve category-level recognition robustness. To support research on asymmetric cross-modal FGVC, we further construct BirdPro, a new bird-oriented audio-visual benchmark, since existing datasets often lack large-scale category-level audio-video associations under non-strict temporal and instance correspondence. BirdPro contains 1,919 audio recordings and 11,965 videos covering 194 bird species. Extensive experiments show that ACF-Net achieves the best results compared with representative baseline methods, outperforming the strongest baselines by 2.97% and 1.92% in the fused and mismatched settings, respectively.

CommentsAccepted by the 9th Chinese Conference on Pattern Recognition and Computer Vision (PRCV 2026). 15 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑