发表机构
Institute of Mathematics and Statistics, University of São Paulo; bluecore(圣保罗大学数学与统计学院; 蓝芯科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究设计并验证了面向语音生物标志物研究的浏览器音频捕获系统VocalCap,其可追溯捕获过程且经测试验证了软件行为,为远程语音研究提供可靠音频数据支持。
AI 中文摘要
远程语音研究通常仅保留最终音频文件,却缺乏关于其捕获、传输、处理及接收过程的完整证据。本文提出VocalCap,一种由机构管控、基于浏览器的系统,可供无技术背景的参与者自主完成语音及相关声学信号的捕获。该系统采用带版本控制的协议驱动工作流程,每份被接收的录音均保留三类关联内容:浏览器原生对象、源自同一MediaStream的客户端无损Float32 WAV文件,以及服务器规范的单声道PCM16 WAV文件,同时关联捕获执行、技术质量、字节级完整性、恢复及转换来源的证据。IndexedDB会保存已被接收的浏览器工件,直至服务器确认,而会话完成需成功验证所有任务及工件。软件测试通过畸形或修改对象、零中断、通道拓扑变体、中断或重复操作等方式对捕获契约发起挑战。对39份经知情同意的试点录音的事后技术审计发现,其中25份为样本完全相同的立体声文件,14份的信号仅局限于左声道。拓扑感知的活动通道选择将这14份受影响文件的规范均方根电平差限制在0.001 dB以内;若采用等权重立体声平均,则会引入约6.02 dB的衰减。生产端到端验证在Chromium和WebKit中完成了两个五任务配置,产生10份被接收录音及30份保留工件,均通过服务器端完整性和格式检查。结果验证了VocalCap在被测浏览器引擎条件下的软件行为,而设备级声学一致性、目标人群可用性、临床有效性及生物标志物性能仍需单独研究。
英文摘要
Remote voice studies often retain a final audio file with limited evidence about how it was captured, transferred, processed, and accepted. This paper presents VocalCap, an institution-controlled, browser-based system for self-guided capture of voice and related acoustic signals by participants without technical training. A versioned protocol drives the workflow. Each accepted recording retains a browser-native object, a client-lossless Float32 WAV derived from the same MediaStream, and a server-canonical mono PCM16 WAV, linked to evidence of capture execution, technical quality, byte-level integrity, recovery, and transformation provenance. IndexedDB preserves accepted browser artifacts until server confirmation, while session completion requires successful verification of every task and artifact. Software tests challenged the acquisition contracts with malformed or altered objects, exact-zero interruptions, channel-topology variants, and interrupted or repeated operations. A post hoc technical audit of 39 consented pilot recordings found 25 sample-identical stereo files and 14 files with signal confined to the left channel. Topology-aware active-channel selection limited the canonical root-mean-square level difference to less than 0.001 dB in all 14 affected files; equal-weight stereo averaging would have introduced approximately 6.02 dB of attenuation. Production end-to-end verification completed two five-task profiles in Chromium and WebKit, yielding 10 accepted recordings and 30 retained artifacts that passed server-side integrity and format checks. The results verify VocalCap's software behavior under the tested browser-engine conditions. Device-level acoustic agreement, target-population usability, clinical validity, and biomarker performance remain subjects for separate studies.