发表机构
IIIT-Delhi(德里印度信息技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出名为Always Alert的智能手机音频遇险检测系统,采用两阶段SVM监督学习框架,经多环境音频训练后,可在日常环境中实现高遇险检测率且误报率低,平均开销仅约每3-4小时1条Facebook帖子。
AI 中文摘要
我们研究了一种不引人注意的全天候(24×7)人类遇险检测与信号系统“Always Alert”,该系统要求智能手机而非其人类所有者保持警戒状态。该系统利用麦克风传感器(每部手机至少配备一个),并假设存在数据网络。我们提出了一种新颖的两阶段监督学习框架,使用支持向量机(SVM),该框架在用户智能手机上执行,用于监测人类处于危险时的自然声音表达——本研究中为尖叫和哭泣。面临的挑战是在普通智能手机用户日常活动时,实现高遇险检测率同时确保误报率处于可管理的开销范围内。我们使用精心选择的遇险音频指纹和各种环境上下文对学习框架进行训练,通过调整音频来优化学习框架,以获得理想的遇险检测率和误报率(FAR)。我们证明了所提出框架在相当具有挑战性的音频环境中检测遇险的能力,进一步利用误报的时间连续性可降低误报率。我们通过志愿者在日常活动中用智能手机录制的数小时音频指纹进行测试,展示了该框架随时随地使用的可行性,我们能够实现高遇险检测率,平均开销约为每3至4小时1条Facebook帖子的量级。
英文摘要
We investigate an unobtrusive and $24\times7$ human distress detection and signaling system, Always Alert, that requires the smartphone, and not its human owner, to be on alert. The system leverages the microphone sensor, at least one of which is available on every phone, and assumes the availability of a data network. We propose a novel two-stage supervised learning framework, using support vector machines (SVMs), that executes on a user's smartphone and monitors natural vocal expressions of fear---screaming and crying in our study---when a human being is in harm's way. The challenge is to achieve a high distress detection rate while ensuring that the false alarm rate is a manageable overhead, while a typical smartphone user goes about living life as usual. We train the learning framework with carefully selected audio fingerprints of distress and of varied environmental contexts. The audio is used to tune the learning framework to obtain a desirable distress detection rate and false alarm rate (FAR). The ability of the proposed framework to detect distress in rather challenging audio environments is demonstrated. Exploiting the time contiguous nature of false alarms further allows us to reduce the FAR. We show the feasibility of using our framework anytime and anywhere by testing it over many hours of audio fingerprints recorded by volunteers on their smartphones, as they went about their daily routines. We are able to achieve high distress detection rates at an average overhead that is equivalent to about 1 facebook post every 3 to 4 hours.