Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Synthetic data uses the distribution(s) of the underlying dataset(s) to generate a totally new dataset that in theory has the same statistical properties of the original dataset. The rub is in making sure that you're actually synthesizing a valid statistical representation of the original dataset, including joint distributions. Otherwise, you wind up with models that won't generalize back to the original data.


To the extent that the synthetic data comes from a generative model, why not sell the parameters of the generative model?


It's being asked a lot here. The answer is that this "generative model" is a utopia




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: