基于生成式动画与生物物理嵌入的rPPG面部视频合成算法
DOI:
CSTR:
作者:
作者单位:

1.杭州电子科技大学计算机学院杭州310018;2.全省(浙江)脑机协同智能技术及应用重点实验室杭州310018

作者简介:

通讯作者:

中图分类号:

TN911.73;TP391

基金项目:

国家自然科学基金面上项目(62573171)、浙江省自然科学基金(LZ25F030005,LY24F020015)、浙江省属高校基本科研业务费专项资金(GK269910299001-009)、 全省(浙江)脑机协同智能技术及应用(含中央引导地方科技发展资金项目编号:2025ZY01045、2026ZY01008;计划编号:2025E10015)项目资助


rPPG facial video synthesis via generative animation and biophysical embedding
Author:
Affiliation:

1.School of Computer Science, Hangzhou Dianzi University, Hangzhou 310018, China; 2.Zhejiang Provincial Key Laboratory of Brain Computer Collaborative Intelligence Technology and Applications, Hangzhou 310018, China

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    远程光电容积脉搏波描记法作为一种非接触式生理监测技术,在临床监护与人机交互领域具有重要价值。然而,深度学习模型在rPPG任务中的性能高度依赖于大规模且多样化的标注数据。现有的基准数据集(如PURE和UBFC-rPPG)受限于受试者规模且运动场景较为单一,导致模型在面对复杂非线性运动干扰时泛化能力不足。为此,提出一种基于生成式动画与生物物理嵌入的面部视频合成框架。该框架采用运动生成-信号嵌入的二阶段策略。首先,利用生成式框架 LivePortrait将驱动视频中的大幅度动作与表情迁移至静态图像,构建高保真的动态面部视频,有效模拟了真实场景中的宏观位移噪声;随后,提出一种基于层分离的生物物理嵌入模型,将生理信号植入视频中。该模型利用高斯大核算子将图像分解为基准层与细节层,在低频分量中引入乘性调制,以更符合皮肤光学特性和血流变化的真实表现,在保留皮肤细节纹理的同时,确保了微弱生理信号在光谱与时域维度的物理准确性。在UBFC-rPPG和PURE数据集上的实验结果表明,在样本受限的条件下,加入本研究生成的合成数据后,PhysNet在UBFC-rPPG数据集上的平均绝对误差(MAE)由1.11 bpm降至0.78 bpm,信噪比(SNR)由5.50 dB提升至5.99 dB;在PURE数据集上的MAE由1.16 bpm降至0.99 bpm,SNR由8.17 dB提升至8.36 dB。结果表明,本研究生成的合成数据能够有效提升rPPG模型的心率估计精度与信号质量,为解决rPPG领域的数据稀缺问题提供了一种兼具物理确定性与运动真实感的可行路径。

    Abstract:

    Remote photoplethysmography (rPPG), as a non-contact physiological monitoring technique, is of great value in clinical monitoring and human-computer interaction. However, the performance of deep learning models in rPPG tasks heavily relies on large-scale and diverse annotated data. Existing benchmark datasets, such as PURE and UBFC-rPPG, are limited by subject scale and relatively simple motion scenarios, which leads to insufficient generalization when models face complex nonlinear motion interference. Therefore, this paper proposes a facial video synthesis framework based on generative animation and biophysical embedding. The framework adopts a two-stage strategy of motion generation and signal embedding. First, the generative framework LivePortrait is used to transfer large-scale motions and facial expressions from driving videos to static images, constructing high-fidelity dynamic facial videos and effectively simulating macroscopic displacement noise in real scenes. Then, a layer-separated biophysical embedding model is proposed to implant physiological signals into videos. This model uses a large-kernel Gaussian operator to decompose the image into a base layer and a detail layer, and introduces multiplicative modulation into the low-frequency component, which better conforms to skin optical characteristics and blood-flow variations. While preserving detailed skin texture, it ensures the physical accuracy of weak physiological signals in both spectral and temporal dimensions. Experimental results on the UBFC-rPPG and PURE datasets show that, under sample-limited conditions, after introducing the synthetic data generated by this paper, the MAE of PhysNet on the UBFC-rPPG dataset decreases from 1.11 bpm to 0.78 bpm, and the SNR increases from 5.50 dB to 5.99 dB; on the PURE dataset, the MAE decreases from 1.16 bpm to 0.99 bpm, and the SNR increases from 8.17 dB to 8.36 dB. The results demonstrate that the synthetic data generated by this paper can effectively improve the heart-rate estimation accuracy and signal quality of rPPG models. This study provides a feasible approach with both physical determinacy and motion realism for addressing data scarcity in the rPPG field.

    参考文献
    相似文献
    引证文献
引用本文

李沂杭,程世超,赵昶辰,张建海.基于生成式动画与生物物理嵌入的rPPG面部视频合成算法[J].电子测量与仪器学报,2026,40(7):44-52

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2026-09-20
  • 出版日期:
文章二维码
×
《电子测量与仪器学报》
关于防范虚假编辑部邮件的郑重公告