It’s not that complicated. Recreating training data via distillation is basically asking structured questions and recording the responses and reformatting that to use as cleaned “good” training data. Much less energy and compute intensive than creating the training data on your own.
I think I remember reading somewhere how Chinese research’s do this by basically using bots and spreading out the distillation to many source queries.