arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

主动视觉语义:用于理解动作中视觉智能的大规模MEG和眼动追踪数据集

Active Visual Semantics: A large-scale MEG and eye-tracking dataset for understanding visual intelligence in action

Philip Sulewski, Carmen Amme, Peter König, Martin N. Hebart, Tim C. Kietzmann

arXiv 2609.01055首次发表:更新:

发表机构

Osnabrück University; Max Planck Institute for Human Cognitive and Brain Sciences; Max Planck School of Cognition; Justus Liebig University Giessen; Center for Mind, Brain and Behavior; University Medical Center Hamburg-Eppendorf(奥斯纳布吕克大学; 马克斯·普朗克人类认知与脑科学研究所; 马克斯·普朗克认知学院; 吉森尤斯图斯·李比希大学; 心智、大脑与行为中心; 汉堡-埃彭多夫大学医学中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发布主动视觉语义(AVS)大规模数据集,包含MEG、眼动追踪等数据,通过ANN编码模型等分析,为探究主动视觉等神经机制提供资源。

AI 中文摘要

本文介绍主动视觉语义(Active Visual Semantics, AVS)数据集,这是一个大规模的脑磁图(magnetoencephalography, MEG)与眼动追踪数据集合,采集自5名参与者在10个会话中自由探索4080张自然场景图像时的相关数据,总计产生超过200000个注视时段。与现有依赖强制中央注视的被动观看模式的神经成像数据集不同,AVS采集的是主动场景探索过程中的脑活动,包括自主产生的扫视运动与注视行为。针对25%的试次开展的语义字幕任务提供了将眼动与场景理解及记忆关联起来的行为学测量指标。除神经与行为数据外,AVS还包含每个注视点对应的物体类别标签、参与者对字幕任务中注视目标外观的人工标注,以及瞳孔动态数据。MEG数据采集期间使用了个体头部固定装置,结合结构磁共振成像(structural MRI)扫描,实现了精准的跨会话源重建。通过人工神经网络(artificial neural network, ANN)编码模型,我们证明尽管主动场景观看对MEG信号质量提出了挑战,但与注视点对齐的单个MEG时段仍包含视觉内容特异性信号。此外,我们采用与注视点对齐的表征相似性分析(representation similarity analysis, RSA),证明可推导得到注视物体类别的平均值,其产生的表征几何结构在不同参与者间具有高度可靠性。在MEG传感器空间与源空间中,该结构均通过其与ANN物体级表征几何结构的稳健对齐得到验证。综上,AVS为研究自然观看过程中主动视觉、物体识别与场景字幕的神经机制,以及眼动行为与记忆编码之间的关系提供了丰富的资源。

英文摘要

Here we present the Active Visual Semantics (AVS) dataset, a large-scale collection of magnetoencephalography (MEG) and eye-tracking data recorded while five participants freely explored 4,080 natural scenes over 10 sessions each, yielding more than 200,000 fixation epochs in total. Unlike existing neuroimaging datasets that rely on passive viewing with enforced central fixation, AVS captures brain activity during active scene exploration, including self-generated saccades and fixations. A semantic captioning task on 25% of the trials provides behavioural measures linking gaze to scene understanding and memory. In addition to neural and behavioural data, AVS includes per-fixation object category labels, human-rated annotations of the appearance of fixation targets in the scene captioning task and pupil dynamics. Individual head stabilisation casts were used during MEG data collection, which alongside with structural MRI scans, enabled precise cross-session source reconstruction. Using artificial neural network (ANN) encoding models we demonstrate that individual fixation-aligned MEG epochs hold visual content-specific signal, despite the challenges that active scene viewing poses for MEG signal quality. Further, we use fixation-aligned representation similarity analysis (RSA) and demonstrate that we can derive fixation object category averages that yield representational geometries which are highly reliable across participants. Both in MEG sensor and source space this structure is validated by its robust alignment with ANN object-level representational geometry. Taken together, AVS provides a rich resource for investigating a large variety of questions regarding the neural mechanisms of active vision, object recognition and scene captioning during natural viewing, and the relationship between gaze behaviour and memory encoding.

Comments21 pages, 3 figures. High-resolution figures are provided at the end of the manuscript PDF. M.N. Hebart and T.C. Kietzmann contributed equally. Corresponding authors: P. Sulewski (phsulewski@gmail.com), T.C. Kietzmann (tim.kietzmann@uni-osnabrueck.de)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑