arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

抹香鲸(Physeter macrocephalus)声音较量中重叠点击序列的视听日记化:基于三水听器阵列

Audiovisual diarization of overlapping click trains in sperm whale (Physeter macrocephalus) vocal sparring using a three-hydrophone array

Lara Berkenbaum, Hervé Glotin, François Sarano, Walter M X Zimmer, Véronique Sarano, Olivier Adam, Pascale Giraudet

arXiv 2609.12593首次发表:更新:

发表机构

Univ Toulon, Aix Marseille Univ; CNRS, LIS; DYNI NATAL; Longitude 181; Jean Le Rond d’Alembert Inst., Sorbonne University, CNRS; Neurosciences Paris-Saclay Inst., Paris-Saclay Univ., CNRS; Univ Toulon, CIAN(土伦大学,艾克斯-马赛大学; 法国国家科学研究中心,信息系统实验室; DYNI NATAL; Longitude 181; 让·勒朗·达朗贝尔研究所,索邦大学,法国国家科学研究中心; 巴黎-萨克雷神经科学研究所,巴黎-萨克雷大学,法国国家科学研究中心; 土伦大学,CIAN)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对抹香鲸声音较量中重叠点击序列的个体归属难题,提出结合三水听器阵列与视频的视听融合方法,实现高精度日记化并成功分配882个点击至特定个体。

AI 中文摘要

在抹香鲸(Physeter macrocephalus)水面互动期间,对其发声进行个体归属因声学重叠、多径传播、身体遮蔽以及相似体型个体间不可区分的脉冲间隔而构成方法论挑战。本文提出了一种多模态视听工作流程,利用便携式三水听器阵列与同步视频相结合,对未成熟雄性在“声音较量”期间产生的未表征点击序列进行去交织和归属。该方法将基于点击间隔动态和常数Q变换的频谱-时间跟踪,与投影到图像平面上的空间到达时间差建模相结合,以进行光学验证。对2,655个手动验证点击的分析表明,纯声学聚类日记化误差保持较低,仅在极端时间叠加时峰值达到25.79%。整合光学模态解决了残余的空间不确定性;尽管近平面阵列几何结构引起垂直模糊性,该系统对主要发声者实现了高达100%的水平视觉一致性。关键的是,该框架成功重建并分配了11个不同且交织的点击序列,总计882个点击,即使在近场触觉约束下也能归因于特定的焦点个体。由于在没有识别发声者的情况下,行为学描述仍不完整,将这些发射与物理运动学明确关联,为定义这种社会声学行为提供了所需的精细尺度分辨率。

英文摘要

The individual attribution of sperm whale (Physeter macrocephalus) vocalizations during surface interactions constitutes a methodological challenge due to acoustic overlaps, multipath propagation, body shadowing, and indistinguishable inter-pulse intervals among similar-sized individuals. A multimodal audio-visual workflow is presented to deinterleave and attribute uncharacterized click trains produced by immature males during ``vocal sparring'' using a portable three-hydrophone array coupled with synchronized video. The approach combines spectro-temporal tracking, based on inter-click interval dynamics and the Constant-Q Transform, with spatial time-difference-of-arrival modeling projected onto the image plane for optical validation. Analysis of 2,655 manually validated clicks shows that purely acoustic clustering diarization errors remain low, peaking at 25.79% only during extreme temporal superpositions. Integrating the optical modality resolves residual spatial indeterminacies; although near-planar array geometry induces vertical ambiguities, the system achieves up to 100% horizontal visual concordance for primary emitters. Crucially, this framework successfully reconstructed and assigned 11 distinct, intertwined click trains totaling 882 clicks to specific focal individuals despite near-field tactile constraints. Because ethological descriptions remain incomplete without identifying the emitter, explicitly correlating these emissions with physical kinematics provides the fine-scale resolution required to define this socio-acoustic behavior.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑