Synthetic data uses the distribution(s) of the underlying dataset(s) to generate a totally new dataset that in theory has the same statistical properties of the original dataset. The rub is in making sure that you're actually synthesizing a valid statistical representation of the original dataset, including joint distributions. Otherwise, you wind up with models that won't generalize back to the original data.