发表机构
The Autonomous Navigation and Sensor Fusion Lab, the Hatter Department of Marine Technologies, University of Haifa(海法大学自主导航与传感器融合实验室,海洋技术哈特学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对深度学习在惯性传感器分类任务中依赖大规模数据集的瓶颈,通过系统实证评估学习曲线收敛率,提出统一框架和经验公式,引入定量稳定点指标,提供可推广框架,优化数据效率,为惯性传感应用记录活动规划提供指南。
AI 中文摘要
深度学习模型对大规模惯性数据集的依赖在基于惯性传感器的分类任务中构成了重大瓶颈,如人类活动识别和智能手机位置识别。这些领域的数据收集复杂、耗时且难以扩展。目前不存在确定达到所需精度水平所需最小样本量的数据驱动指南。本研究对惯性分类中的学习曲线收敛率进行了系统实证评估。引入统一框架分析二分类和多分类场景下的分类性能,推导估计性能与数据集大小关系的经验公式。在六个总计102.7小时惯性测量的真实数据集上测试表明,无论任务复杂性如何,准确率都遵循一致的对数增长模式。据此提出定量稳定点指标,分析表明模型达到实际稳定所需样本比传统启发式方法建议的少得多。最终提供可推广框架从小规模试点研究推断总数据需求,优化记录工作与模型可靠性之间的权衡。这些发现将主流范式从最大化数据量转向优化数据效率,为惯性传感应用中的记录活动规划提供具体的、数据支持的指南。
英文摘要
Deep learning models dependency on large-scale inertial datasets presents a significant bottleneck in inertial sensor-based classification tasks, such as human activity recognition and smartphone location recognition. In these domains, data collection requires massive recording campaigns that are complex, time-consuming, and difficult to scale. Currently, data-driven guidelines for determining the minimum sample size required to reach a desired accuracy level do not exist. To address this gap, this study presents a systematic empirical evaluation of learning curve convergence rates in inertial classification. We introduce a unified framework that analyzes classification performance under both binary and multi-class scenarios, and derive an empirical formula to estimate performance relative to dataset size. Testing across six diverse, real-world datasets totaling 102.7 hours of inertial measurements demonstrates that accuracy follows a consistent logarithmic growth pattern, regardless of task complexity. Leveraging this finding, we propose a quantitative stability point metric, defined as the sample size required for the learning curve to stabilize within a predefined mean absolute percentage deviation of its asymptotic maximum. Our analysis reveals that models often reach practical stability with substantially fewer samples than traditional heuristics suggest. Ultimately, we offer a generalizable framework to extrapolate total data requirements from small-scale pilot studies, optimizing the tradeoff between recording effort and model reliability. These findings shift the prevailing paradigm from maximizing data volume toward optimizing data efficiency, offering concrete, data-backed guidelines for planning recording campaigns in inertial sensing applications.
Comments1 pages, 17 figures, 15 tables