arXiv cs.LGOctober 2, 2026
Training-Aware Target Coverage for Synthetic Data Selection
Excerpt
arXiv:2610.00814v1 Announce Type: new Abstract: Synthetic data are increasingly used to scale LLM training, yet more synthetic data do not necessarily produce better models. Useful synthetic data must add information relevant to the target task without introducing errors that offset their benefit, and the value of an example can change as the training set grows. We develop a linear theory that characterizes this tradeoff and determines where synthetic data are useful, how much should be added, a