← Back to all articles
arXiv cs.LGOctober 1, 2026

Synthetic Data Characterization via Training Dynamics

Excerpt

arXiv:2609.39447v1 Announce Type: cross Abstract: Interpreting properties of LLM-generated data is important for understanding its utility and limitations across learning tasks. In this work, we characterize synthetic data through sample-level learnability, studying variation among LLM families and scales, alongside human-written data as a reference. We first generate synthetic datasets spanning single- and multi-label classification, labeling, and tree prediction tasks. We then derive empirical