

It’s not that complicated. Recreating training data via distillation is basically asking structured questions and recording the responses and reformatting that to use as cleaned “good” training data. Much less energy and compute intensive than creating the training data on your own.
I think I remeber reading somewhere how Chinese research’s do this by basically using bots and spreading out the distillation to many source queries.



Yeah I know they’re not the only companies doing distillation, it’s just currently in the news and on peoples minds.