arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

隐空间中的 galvanic 前庭刺激(GVS)

Galvanic Vestibular Stimulation in Latent Space

Zhi Liu, Tatsuki Fushimi, Yoichi Ochiai

arXiv 2607.26659首次发表:更新:

AI 中文总结

该研究针对 GVS 波形与目标状态匹配合成的挑战,构建关联数据集与检索引导生成模型,验证了文本条件下 GVS 合成的可行性,为将 GVS 开发为具身反馈可编程模态提供支持。

AI 中文摘要

galvanic 前庭刺激(GVS)被广泛用于调节自我定向、平衡与运动感知;频率编码线索的可区分性进一步表明,其作为具身反馈独立模态的潜力。然而,合成与目标事件或身体状态一致的 GVS 波形仍具挑战性。GVS 波形结合了电流方向、强度、持续时间以及起止过渡,但这些参数如何共同塑造用户的感知与关联反应仍未得到充分探索。为解决这一空白,本文贡献了一个将 GVS 波形与自由形式体验描述关联的数据集,以及一个用于从目标描述合成候选波形的检索引导生成模型。该数据集包含 100 个 GVS 波形和从 16 名参与者处收集的 1526 条有效自由形式感觉描述。语义分析揭示了多样的运动与力相关感觉、局部身体感觉以及情境关联。与保留参与者的置换基线相比,同一波形引发的描述覆盖更少的语义类别(8.18 对 9.45),且表现出更高的主导类别比例(26.97% 对 21.25%;两者 P < 0.001)。基于该数据集,本文将生成模型实现为检索引导的一维卷积变分自编码器。一项独立行为研究招募了 10 名未参与数据集收集的参与者,其区分一致与不一致的波形-视觉线索配对的表现显著高于随机水平,准确率为 63.33%,d-prime = 0.70,且 p < 0.001。综上,这些发现证明了文本条件下 GVS 合成的可行性,并支持将 GVS 开发为交互式场景中语义一致的具身反馈可编程模态。

英文摘要

Galvanic vestibular stimulation (GVS) is widely used to modulate self-orientation, balance, and motion perception; the discriminability of frequency-encoded cues further suggests its potential as a standalone modality for embodied feedback. However, synthesizing GVS waveforms congruent with target events or bodily states remains challenging. GVS waveforms combine current direction, intensity, duration, and onset and offset transitions, yet how these parameters jointly shape users' perceptual and associative responses remains underexplored. To address this gap, we contribute a dataset linking GVS waveforms to free-form experience descriptions, as well as a retrieval-guided generative model for synthesizing candidate waveforms from target descriptions. The dataset comprises 100 GVS waveforms and 1,526 valid free-form sensation descriptions collected from 16 participants. Semantic analysis revealed diverse motion- and force-related sensations, localized bodily sensations, and situational associations. Compared with a participant-preserving permutation baseline, descriptions elicited by the same waveform covered fewer semantic categories (8.18 vs. 9.45) and exhibited a higher dominant-category proportion (26.97% vs. 21.25%; both P < 0.001). Building on this dataset, we implemented the generative model as a retrieval-guided one-dimensional convolutional variational autoencoder. An independent behavioral study recruited 10 participants who had not contributed to the dataset collection. Performance in discriminating congruent from incongruent waveform-visual cue pairings was significantly above chance, with an accuracy of 63.33%, d-prime = 0.70, and p < 0.001. Together, these findings demonstrate the feasibility of text-conditioned GVS synthesis and support the development of GVS as a programmable modality for semantically congruent embodied feedback across interactive scenarios.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑