arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PhysVR:面向远程生理测量的、由视觉语言模型引导的感知干扰的时间特征优化框架

PhysVR: Vision-Language Model Guided Interference-aware Temporal Feature Refinement for Remote Physiological Measurement

Zixu Li, Jianjun Qian, Hang Shao, Daoheng Li, Lei Luo, Jian Yang

arXiv 2608.29663首次发表:更新:

发表机构

School of Computer Science and Engineering, Nanjing University of Science and Technology; College of Computer Science and Technology, Qingdao University(南京理工大学计算机科学与工程学院; 青岛大学计算机科学技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对远程生理测量中rPPG易受干扰的问题,提出PhysVR框架,通过视觉语言模型引导优化时间特征,在5个公共基准上表现优于现有方法。

AI 中文摘要

远程光电容积描记法(rPPG)可从面部视频中实现非接触式生理测量,但其与脉搏相关的细微变化易受光照变化、头部运动、面部模糊及感兴趣区域不稳定的影响。现有方法主要在特征学习阶段抑制干扰,却极少探究学习到的时间特征是否仍受干扰影响,以及在rPPG估计前如何进一步抑制此类干扰。为解决该局限,本文提出PhysVR,一种由视觉语言模型(VLM)引导的、感知干扰的rPPG估计时间特征优化框架。具体而言,生理主干网络生成全局时间特征及粗略rPPG预测,基于此从局部时间特性构建信号衍生的生理可靠性证据;并行地,冻结的视觉语言模型在面向干扰的提示下处理采样的面部帧,证据头从VLM输出中提取视觉干扰证据;时间交叉注意力将生理证据、视觉证据与全局时间特征融合,构建感知干扰的时间上下文。在该上下文引导下,共享时间修正单元执行通用优化,同时4个干扰特定专家通过自适应路由选择性抑制不同干扰;优化后的时间特征随后用于最终rPPG估计。在5个公共基准上的大量实验表明,PhysVR在数据集内及跨数据集评估协议下均持续优于代表性方法。

英文摘要

Remote photoplethysmography (rPPG) enables contactless physiological measurement from facial videos, yet its subtle pulse-related variations are easily affected by illumination variation, head motion, facial blur, and region-of-interest instability. Existing methods mainly suppress interference during feature learning, while whether the learned temporal features remain affected by interference and how to further suppress such interference before rPPG estimation are rarely examined. To address this limitation, we propose PhysVR, a vision-language model guided interference-aware temporal feature refinement framework for rPPG estimation. Specifically, a physiological backbone produces global temporal features and a coarse rPPG prediction, from which signal-derived physiological reliability evidence is constructed from local temporal characteristics. In parallel, a frozen vision-language model processes sampled facial frames under an interference-oriented prompt, and an evidence head extracts visual interference evidence from the VLM output. Temporal cross-attention integrates the physiological and visual evidence with the global temporal features to construct interference-aware temporal context. Guided by this context, a shared temporal correction unit performs general refinement, while four interference-specific experts selectively suppress different interference through adaptive routing. The refined temporal features are then used for final rPPG estimation. Extensive experiments on five public benchmarks demonstrate that PhysVR consistently outperforms representative methods under both intra-dataset and cross-dataset evaluation protocols.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑