Skip to content
Case Study

When You Can't Use Real Patient Data: Synthetic Tabular Data with Diffusion Models

Facts checked against primary sources on 2026-09-27.

In clinical research, data is almost always the bottleneck. The real data exists, but it is locked behind IRB approvals, HIPAA compliance requirements, data use agreements, and legal review that can take months. Meanwhile, the ML team needs data to develop and validate models.

On a clinical research engagement between late 2024 and early 2026, we addressed this by generating high-fidelity synthetic tabular data using diffusion models. Below is how that worked and how we judged whether the output was usable.

Synthetic data can still leak

Synthetic data is often assumed to be safe to share, but data that stays too faithful to the source can still leak patient information, through membership inference or attribute inference attacks. Stadler, Oprisanu and Troncoso (USENIX Security 2022) evaluated a range of generative models and found that synthetic data either failed to prevent these inference attacks or lost its utility. A synthetic dataset where every row is basically a slightly-perturbed real patient isn't private at all, whatever it's labeled as.

Under HIPAA, data counts as de-identified only through one of the two routes in 45 CFR 164.514(b): Safe Harbor (removing 18 listed identifier types, with no actual knowledge that what remains could identify someone) or Expert Determination (a qualified expert documents that the risk of re-identification is "very small"). Nothing in the rule treats model-generated records as de-identified by default. Synthetic data derived from PHI still has to clear one of those bars, and in practice that usually means an expert determination backed by privacy testing like the kind described below.

The generator therefore has to keep the feature distributions, correlations, and conditional relationships that downstream models rely on, without letting anyone reconstruct or infer anything about an individual real patient. Stronger privacy protection usually lowers fidelity, so the acceptable trade-off is set per project by the intended use and the sensitivity of the source data.

Why not just use a GAN

CTGAN and TVAE, both introduced in Xu et al., "Modeling Tabular Data using Conditional GAN" (NeurIPS 2019), are the usual starting point for synthetic tabular data. They were designed for mixed discrete and continuous columns, multimodal distributions, and imbalanced categories, so they are a reasonable baseline. Clinical data still pushes them hard: it mixes continuous lab values with binary flags, ordinal severity scores, and categorical diagnoses in the same table, it has missingness patterns that carry real clinical meaning rather than being random noise, and it has strong conditional structure (a patient with condition A will almost always show lab value B in a specific range). GANs are trained as a min-max game between generator and discriminator, which is known to be prone to unstable training and mode collapse, and in practice the learned distribution can drift away from the real conditional relationships.

A diffusion model avoids the adversarial game: it is trained with a simple denoising objective and learns the joint distribution by reversing a gradual noising process. The best-known tabular version, TabDDPM (Kotelnikov et al., ICML 2023), combines Gaussian diffusion for numerical columns with multinomial diffusion for categorical ones and outperformed GAN and VAE baselines, including CTGAN and TVAE, across its benchmark datasets. That approach held up much better against clinical data's quirks in practice. We used a score-based diffusion model adapted for mixed-type tabular data, with separate noise schedules for continuous and categorical columns.

Keep the missingness pattern

Don't impute missingness away before training. Whether a lab test was even ordered tells you something about the patient's condition. A model trained on pre-imputed data throws that signal out. In one large study, Agniel, Kohane and Weber (BMJ, 2018) found that across 272 common lab tests, the timing and frequency of ordering were often more predictive of three-year survival than the test results themselves. We added a binary missingness indicator for every column with more than 5% missing values and trained the diffusion model to jointly generate the value and the missingness indicator together, so the not-missing-at-random structure survives into the synthetic data.

Evaluating fidelity and privacy

"It looks similar" isn't a metric, so we leaned on three checks that each catch a different failure mode. First, distributional fidelity: Kolmogorov-Smirnov tests per continuous feature, chi-squared for categoricals, and the Frobenius norm between real and synthetic correlation matrices to catch relationships a marginal-only check would miss. Second, train-on-synthetic-test-on-real (TSTR, a protocol proposed by Esteban, Hyland and Rätsch, 2017): train a downstream classifier purely on synthetic data, evaluate it against a held-out real test set, and compare to the same classifier trained on real data. The AUC gap is the actual fidelity loss that matters. Ours had to land within 0.03 of the real-data AUC, and the confusion-matrix structure had to look similar too, not just the headline number. Third, a privacy audit: a membership-inference attack using the shadow-model technique from Shokri et al. (IEEE S&P 2017), whose accuracy (or AUC) should land close to 0.5, random chance, if the attacker cannot tell members from non-members, plus a nearest-neighbor distance ratio check: if synthetic records sit suspiciously close to real ones in feature space, the model is memorizing rather than generating.

One caveat on the privacy audit: a passing attack result only shows that the specific attacks you ran did not succeed. A provable bound requires a mechanism such as differential privacy, which costs fidelity. The audit is useful input to an expert determination, but it does not by itself establish that the data is anonymous.

None of this means anything to a clinical researcher

KS tests and NNDR scores mean nothing to clinical researchers or a compliance reviewer, and shouldn't need to. What they actually wanted to know boiled down to two questions: does the synthetic data lead to the same clinical conclusions as the real data, and can a patient be identified from it. So the evaluation got translated into two slides instead of a metrics table: one overlaying real and synthetic survival curves (nearly identical curves were what actually built trust, not a p-value), and one showing that the nearest synthetic record to any real patient stayed outside a defined clinical threshold on the variables that mattered. That second slide was the privacy assurance they needed, in a form they could actually evaluate themselves.

Their pushback on that framing fed back into which metrics got prioritized in the next iteration: the two-slide version wasn't just a communication layer bolted on after the fact, it changed what we measured.

What took the most time

Model training turned out to be the easy part. Getting missingness right, building an evaluation framework rigorous enough to trust, and getting non-technical stakeholders to actually believe the result took longer than the diffusion model itself did.

Written by Sebastian Cepeda. Senior Machine Learning Engineer specializing in clinical AI.

Need something like this for your team?

I work as a machine learning consultant, production ML systems, MLOps, and clinical/LLM applications.

Get in touch โ†’ ยท See my work