PhysAgent:用于可靠远程心率估计的多智能体框架
PhysAgent: A Multi-Agent Framework for Reliable Remote Heart Rate Estimation
浏览论文内容
中文总结 AI 辅助
PhysAgent是一种推理时多智能体候选验证框架,以多个基础rPPG估计器输出为待验证生理假设,结合Qwen3-VL-4B多模态大语言模型推理与确定性融合,提升远程心率估计的稳定性与可靠性。
中文摘要 AI 辅助
远程光体积描记法(rPPG)可从面部视频实现非接触式心率估计,但其微弱的生理信号易受运动、光照变化、遮挡、皮肤外观差异及设备噪声干扰。现有rPPG方法通常依赖单一模型直接预测心率或恢复脉搏波形,而不同的强估计器可能对同一视频产生相互冲突但各自看似合理的候选结果。为解决这些冲突,我们提出PhysAgent,这是一种推理时的多智能体候选验证框架。与直接预测方法不同,PhysAgent既不训练新的基础rPPG模型,也不要求多模态大语言模型(MLLM)直接输出心率。相反,它将多个基础估计器的输出视为待验证的生理假设,并使用轻量级4B参数的多模态大语言模型Qwen3-VL-4B,针对视频条件、信号可靠性及候选分歧进行多智能体推理。确定性生理验证器会检查融合提案,可复现的数值融合流程则生成最终心率。在多个公开rPPG基准上的实验结果表明,PhysAgent可提升不同数据集及源域设置下的融合稳定性与可靠性,同时避免直接MLLM预测或无约束集成融合存在的不可复现性及生理不一致问题。代码即将发布。
英文摘要
Remote photoplethysmography (rPPG) enables non-contact heart-rate estimation from facial videos, but its weak physiological signal is easily corrupted by motion, illumination changes, occlusion, skin-appearance variation, and device noise. Existing rPPG methods typically rely on a single model to directly predict heart rate or recover pulse waveforms, while different strong estimators may produce conflicting yet individually plausible candidates for the same video. To resolve these conflicts, we propose PhysAgent, an inference-time multi-agent candidate-verification framework. Unlike direct prediction approaches, PhysAgent neither trains a new base rPPG model nor asks Multimodal Large Language Models (MLLMs) to output heart rate directly. In contrast, it treats outputs from multiple base estimators as physiological hypotheses to be verified and uses a lightweight 4B MLLM, Qwen3-VL-4B, to drive multi-agent reasoning over video conditions, signal reliability, and candidate disagreement. A deterministic physiological verifier checks the fusion proposal, and a reproducible numerical fusion process produces the final heart rate. Experimental results on multiple public rPPG benchmarks show that PhysAgent improves fusion stability and reliability across different datasets and source-domain settings, while avoiding the irreproducibility and physiological inconsistency of direct MLLM prediction or unconstrained ensemble fusion. The code will be released soon.