面向鲁棒远程光电容积描记的双分支时空解耦与重构方法
DOI:
CSTR:
作者:
作者单位:

1.杭州电子科技大学计算机学院杭州310018;2.全省(浙江)脑机协同智能技术及应用重点实验室杭州310018

作者简介:

通讯作者:

中图分类号:

TN911.73;TP391

基金项目:

国家自然科学基金面上项目(62573171)、浙江省自然科学基金项目(LZ25F030005,LY24F020015)、浙江省属高校基本科研业务费专项资金(GK269910299001-009)、 全省(浙江)脑机协同智能技术及应用重点实验室(含中央引导地方科技发展资金:2025ZY01045、2026ZY01008;计划编号:2025E10015)项目资助


Dual-branch spatiotemporal disentanglement and reconstruction for robust remote photoplethysmography
Author:
Affiliation:

1.School of Computer Science, Hangzhou Dianzi University, Hangzhou 310018, China; 2.Zhejiang Provincial Key Laboratory of Brain Computer Collaborative Intelligence Technology and Applications, Hangzhou 310018, China

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    远程光电容积描记技术可利用普通摄像设备从人脸视频中恢复脉搏相关生理信号,但面部颜色变化幅度微弱,易受头部运动、姿态变化、光照波动、阴影遮挡和成像噪声等非生理因素干扰。现有深度学习方法多采用端到端回归范式,通常缺乏对脉搏成分与噪声成分生成机制的显式刻画,导致模型中间状态可解释性不足。针对上述问题,提出一种面向鲁棒rPPG的物理先验引导双分支时空解耦与重构网络PNRNet。该方法将面部RGB时序变化建模为生理脉搏分量与非生理噪声分量的叠加,分别通过独立3D卷积分支学习脉搏相关动态和干扰相关动态,并引入特征正交约束降低两类隐表征的冗余耦合。受二色反射模型启发,进一步设计RGB时序重构模块,将预测脉搏信号和噪声信号投影回RGB观测域,使网络同时受到参考PPG波形监督和视频观测域重构约束。实验结果表明,PNRNet在iBVP、UBFC-rPPG和PURE数据集上的平均绝对值(MAE)分别为3.02、1.18和1.78 bpm,其中在iBVP数据集上较最佳对比方法降低88%。波形分析进一步显示,脉搏分支输出与真实PPG在主周期和整体趋势上保持较好一致性,噪声分支输出能够反映部分运动和光照引起的非生理波动。结果说明,结构化引入物理先验有助于提升rPPG模型的鲁棒性和可解释性。

    Abstract:

    Remote photoplethysmography (rPPG) can recover pulse-related physiological signals from facial videos captured by ordinary cameras. However, the cardiac-induced color variations are extremely subtle and are easily corrupted by head motion, pose variation, illumination fluctuation, shadow occlusion, and imaging noise. Most existing deep learning methods formulate rPPG estimation as an end-to-end regression problem, but they usually lack explicit modeling of the generation mechanisms of pulse and noise components, making the intermediate representations difficult to interpret. To address this problem, this paper proposes PNRNet, a physics-guided dual-branch spatiotemporal disentanglement and reconstruction network for robust rPPG. The proposed method models facial RGB temporal variations as a superposition of physiological pulse components and non-physiological noise components. Two independent 3D convolutional branches are used to learn pulse-related and interference-related dynamics, and a feature orthogonality constraint is introduced to reduce redundant coupling between the two latent representations. Inspired by the dichromatic reflection model, an RGB temporal reconstruction module projects the predicted pulse and noise signals back to the RGB observation domain, enabling the network to be jointly supervised by reference PPG waveforms and video-domain reconstruction. Experimental results show that PNRNet achieves MAE of 3.02, 1.18, and 1.78 bpm on the iBVP, UBFC-rPPG, and PURE datasets, respectively, with an 8.8% MAE reduction over the best competing method on iBVP. These results indicate that the proposed pulse-noise disentanglement and RGB observation-domain reconstruction mechanism helps improve the robustness and interpretability of rPPG models under complex interference conditions.

    参考文献
    相似文献
    引证文献
引用本文

张建海,姜丰,程世超,赵昶辰.面向鲁棒远程光电容积描记的双分支时空解耦与重构方法[J].电子测量与仪器学报,2026,40(7):13-23

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2026-09-20
  • 出版日期:
文章二维码
×
《电子测量与仪器学报》
关于防范虚假编辑部邮件的郑重公告