arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36217cs.CV

稀疏视角可解释的三维动物行为表征用于神经编码与解码

Sparse-View Interpretable 3D Animal Behavior Representations for Neural Encoding and Decoding

Xinming Dai, Qihang Jin, Tianshu Tan, Baiyuan Chen, Hanrui Lyu, Lenny Aharon, Kyle Daruwalla, Xun Helen Hou, Matthew R. Whiteway, Liam Paninski, Yizi Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

提出SABLE框架,利用几何归纳偏置从稀疏视角自监督学习可解释的三维动物行为表征,在神经编码解码中达到或超越SOTA,并支持零样本泛化。

中文摘要 AI 辅助

对大脑功能的深入理解需要对行为进行精确、结构化的表征,然而从视频中提取适合科学分析的行为表征仍然是一个基本挑战。许多先前的研究通过姿态估计或非线性视频嵌入来表征行为。然而,姿态追踪会丢弃预定义关键点之外的丰富信息,而非线性视频嵌入缺乏可解释性。我们通过SABLE(稀疏视角动物行为潜在嵌入)来解决这一局限性,这是一种自监督框架,利用几何归纳偏置来学习行为表征。通过用单目深度和姿态估计的先验信息增强多视角Transformer,SABLE能够从极其稀疏的视角重建三维动物行为,同时学习显式的三维潜在结构。在没有真实三维标签的情况下,它能够可靠地从双视角视频中恢复三维行为,而最先进(SOTA)方法则失败或产生退化解。在国际脑实验室和Cheese3D数据集上,我们证明SABLE学习到的三维表征在神经编码和解码方面达到或超过了先前SOTA的性能。一旦在多个动物上预训练完成,SABLE即可作为现成模型,无需针对特定动物的校准或重新训练即可零样本泛化到未见过的动物。我们的方法建立了能够捕捉复杂行为的三维感知视频嵌入,为研究大脑-行为关系开辟了新途径。

英文摘要

A deeper understanding of brain function requires a precise, structured characterization of behavior. Yet, extracting behavioral representations from video in a form suitable for scientific analysis remains a fundamental challenge. Many prior studies represent behavior via pose estimation or nonlinear video embeddings. However, pose tracking discards rich information beyond predefined keypoints, while nonlinear video embeddings lack interpretability. We address this limitation with SABLE (Sparse-view Animal Behavior Latent Embeddings), a self-supervised framework that leverages a geometric inductive bias to learn behavior representations. By augmenting a multi-view transformer with priors from monocular depth and pose estimation, SABLE reconstructs 3D animal behavior from extremely sparse views while learning explicit 3D latent structure. Without ground-truth 3D labels, it reliably recovers 3D behavior from two-view videos, whereas state-of-the-art (SOTA) methods fail or yield degenerate solutions. Across the International Brain Lab and Cheese3D datasets, we demonstrate that SABLE learns 3D representations that match or exceed prior SOTA performance in neural encoding and decoding. Once pretrained across animals, SABLE serves as an off-the-shelf model that generalizes zero-shot to unseen animals without animal-specific calibration or retraining. Our method establishes 3D-aware video embeddings that capture complex behavior, opening new avenues for studying brain-behavior relationships.

发表机构

  • Harvard University(哈佛大学)
  • University of Cambridge(剑桥大学)
  • Northwestern University(西北大学)
  • Cold Spring Harbor Laboratory(冷泉港实验室)
  • Columbia University(哥伦比亚大学)
  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

↑