Evaluate VAEs vs. Diffusion for Synthetic Tabular-Data Generation
Overview
What this challenge is about.
Train a tabular diffusion model on patient data, compare it to a VAE baseline, and write a privacy memo. Finish with a verifiable certificate.
The scenario
The startup (around 35 people, post-clinical-validation) routinely runs joint studies with hospital partners in Israel and Germany; a usable synthetic-data generator removes a 6-month privacy-review bottleneck from each new collaboration.
The Brief
What you'll do, and what you'll demonstrate.
Compare a tabular diffusion model with a VAE baseline on synthetic patient-record generation across fidelity, utility, and privacy.
Earning criteria — what you'll demonstrate
- Train tabular diffusion and VAE generators on real data
- Evaluate synthetic data across fidelity, utility, and privacy
- Run a basic membership-inference attack as privacy evaluation
- Communicate privacy trade-offs to platform leadership
Program Fit
Where this fits in your program.
Sharpens the same skills your degree expects you to demonstrate.
Generative AI
Master · Generative Ai
Strong alignment
This challenge maps to Generative AI at the Master level. It sharpens the same practical skills your coursework expects — but in a real industry context with actual constraints and deliverables.
Skills
Skills you'll demonstrate.
Each one shows up on your verified credential.
- Tabular Diffusion
Apply tabular diffusion to solve real industry problems and demonstrate production-level capability.
- Vae
Apply vae to solve real industry problems and demonstrate production-level capability.
- Synthetic Data
Apply synthetic data to solve real industry problems and demonstrate production-level capability.
- Privacy Evaluation
Apply privacy evaluation to solve real industry problems and demonstrate production-level capability.
- Pytorch
Apply pytorch to solve real industry problems and demonstrate production-level capability.
- Evaluation
Apply evaluation to solve real industry problems and demonstrate production-level capability.
Careers
Career paths this challenge builds toward
Completing this challenge demonstrates skills that transfer directly to these roles:
Research Scientist
Running a tabular-generator comparison with privacy + utility + fidelity evaluation is exactly the day-one work of a research scientist at any healthtech or privacy-AI team.
This challenge sharpens
- tabular-diffusion
- vae
- synthetic-data
AI Safety Researcher
Implementing a membership-inference attack as part of privacy evaluation is core AI safety work in regulated-data settings.
This challenge sharpens
- privacy-evaluation
- evaluation
- synthetic-data
Data Scientist
Comparing two generators on real downstream utility transfers directly to data-science roles where synthetic data unblocks collaboration.
This challenge sharpens
- evaluation
- vae
- pytorch