AI 中文总结
本研究提出基于视觉-语言模型的流程,结合文本线索、视觉特征与已知负载质量,从RGB视频估算手工物料搬运的手力,验证了该方法的可行性及不同相机条件的性能差异。
AI 中文摘要
外部手力是职业体力暴露与损伤风险生物力学分析的重要输入,但手工物料搬运(MMH)过程中的连续力测量通常需要使用带传感器的物体或专用传感设备。本研究评估了一种基于视觉-语言模型(VLM)的流程,该流程结合任务特定的文本线索、视觉表征及已知的箱子质量,从RGB视频估算动态三轴双侧外部手力。35名健康年轻人完成5项MMH任务,包括搬运、携带、推和拉,箱子质量为6、9和12kg。该流程使用文本引导的参与者与被搬运物体感兴趣区域(ROI)定位、预训练视觉Transformer特征提取及基于Transformer的时间回归。通过留一受试者交叉验证,在7种相机视角条件(3种单视角和4种多视角条件)及4种ROI策略下评估性能。总体而言,水平和内外侧力分量的均方根误差约为4.7-5.6N,垂直分量约为10.6-11.0N。将被搬运物体作为第二个ROI通常可改善力估算,在单相机条件下收益最大,而像素级分割几乎未提供额外提升。多相机采集对峰值力估算的益处最明显,尤其对垂直分量,而相机配置间整体帧级误差的差异相对较小。这些发现证明,无需在工人或被搬运物体上放置传感器作为模型输入,即可从RGB视频和已知负载质量估算连续双侧定向手力,支持开发更具可扩展性的职业体力暴露与风险评估方法。
英文摘要
External hand forces are important inputs to biomechanical analyses of occupational physical exposure and injury risk, yet continuous force measurements during manual material handling (MMH) typically requires instrumented objects or specialized sensing. We evaluated a vision-language model (VLM)-based pipeline that combines task-specific textual cues, visual representations, and known box mass to estimate dynamic, triaxial, bilateral external hand forces from RGB video. Thirty-five healthy young adults performed five MMH tasks involving lifting, carrying, pushing, and pulling with box masses of 6, 9, and 12 kg. The pipeline used text-guided localization of participant and handled-object regions of interest (ROIs), pretrained vision-transformer feature extraction, and transformer-based temporal regression. Performance was evaluated using leave-one-subject-out validation across seven camera-view conditions (three single-view and four multi-view conditions) and four ROI strategies. Overall, root mean square error was ~4.7-5.6 N for the horizontal and mediolateral force components and ~10.6-11.0 N for the vertical component. Including the handled object as a second ROI generally improved force estimation, with some of the largest benefits under single-camera conditions, whereas pixel-level segmentation provided little additional improvement. Multi-camera capture provided the clearest benefit for peak-force estimation, particularly for the vertical component, whereas differences in overall frame-level error among camera configurations were comparatively modest. These findings demonstrate the feasibility of estimating continuous, bilateral, directional hand-force estimates from RGB video and known load mass without requiring sensors on the worker or handled objects as model inputs, supporting the development of more scalable occupational physical exposure and risk assessments.