arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从野外拍摄的单张人体图像估计体重与身高

Weight and Height Estimation from a Single Human Image Captured in the Wild

Hira Yaseen, Arif Mahmood, Waqas Sultani

arXiv 2607.26104首次发表:更新:

发表机构

Information Technology University (ITU)(信息技术大学(ITU))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出含6105张带真实标签图像的野外人体数据集,用多模态深度神经网络从单张人体图像估计BMI、体重与身高,实验显示全身图像估计效果优于半身和面部图像。

AI 中文摘要

体重和身高这类人体生理特征是反映身心健康、日常生活习惯及财务状况的重要指标。身体质量指数(BMI)是融合了体重与身高特征的知名衡量指标,被用作自我监测工具,对个人生活有长期影响,例如可帮助预测多种疾病风险、预估寿命。利用野外拍摄的单张人体图像自动估计BMI是一项具有挑战性的任务,因为人体姿态、相机几何、个人外观及干扰背景存在广泛变化。本文探讨了采用单任务与多任务学习的深度神经网络性能,通过运用RGB、深度图、姿态亲和图、边缘图等不同模态,从社交网站上的日常生活图像中预测BMI、体重与身高。目前尚无公开可用的用于BMI估计的全身图像数据集,因此本文提出了一个新数据集,包含6105张带有身高、体重和BMI真实标签的图像。该数据集在野外收集,包含不同种族、年龄组和性别的图像,涵盖正面、背面、全身、半身、侧面姿态,以及背景和尺度各异的镜像自拍,可能存在遮挡部分或全部面部的伪影。使用VGG、Densenet、ResNet等不同CNN骨干,仅对全身、半身和面部图像进行了大量实验,实验结果表明,在野外场景中,全身图像的估计效果优于半身和面部图像。

英文摘要

A person's physical characteristics such as weight and height are important indicators of his physical and mental health, daily life routines and finances. Body Mass Index (BMI) is a well known measure that encodes the characteristics of both the weight and the height. BMI has been used as a self-monitoring tool, and it has long-term implications on one's life. For example, it may help predicting the risk of various diseases and estimating longevity. Automatic BMI estimation using a single person image in the wild is a challenging task due to wide variations in human pose, camera geometry, personal appearance and distracting backgrounds. In this paper, we explore the performance of deep neural networks using single and multi-task learning by employing different modalities including RGB, depth-maps, pose-affinity maps, and edge-maps to predict BMI, weight, and height from daily life images available on social networking websites. Currently, no full body image dataset for BMI estimation is publicly available, therefore we propose a new dataset consisting of 6105 images with ground truth labels of height, weight and BMI. Our proposed dataset is collected in the wild containing images from various ethnicity and distributed over varying age groups and gender. It consists of frontal, back, full and half body, side poses, mirror selfies with varying backgrounds and scale variations and may contain artifacts hiding partial or full face. Extensive experimentation is performed using full body, half body and face images only using different CNN backbones including VGG, Densenet and ResNet. Our experimental results demonstrate that full body images have produced better results than the other half body and facial images in the wild.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑