Evaluation of synthetic data generation for longitudinal 6-year follow-up subclinical cardiovascular disease data using GAN models
Although largely preventable, cardiovascular disease (CVD) is the leading cause of death worldwide. Synthetic data generation (SDG) of health data can facilitate data sharing for research and development of personalised cardiovascular disease prevention tools. Current state-of-the-art SDG methods are limited to cross-sectional data or focused on short-term or in-patient EHR applications lacking…
Cardiovascular disease (CVD) remains the leading cause of death globally. Synthetic data generation (SDG) of health data enables research collaboration and the development of personalized prevention tools. However, existing SDG methods primarily focus on cross-sectional or short-term data, lacking benchmarks for long-term, longitudinal applications. This study assessed the performance of two tabular generative adversarial networks (GANs), TGAN and CTGAN, on a structured 6-year longitudinal trial dataset for CVD prevention.
Both TGAN and CTGAN exhibited better privacy performance compared to a rudimentary perturbation method. CTGAN demonstrated superior fidelity relative to TGAN. While both models maintained comparable privacy levels, CTGAN underperformed in numerical fidelity of continuous variables and overall temporal fidelity compared to the perturbation method. Notably, CTGAN's performance declined for younger age groups compared to older individuals.
This research marks the first evaluation of CTGAN on a longitudinal dataset and suggests that refining CTGAN's handling of temporal characteristics and imbalanced subgroups is necessary to enhance its performance in CVD prevention studies.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.