← Back to all articles
arXiv cs.LGAugust 17, 2026

GALA: Generation-Aware Cross-Modal Alignment for Text-to-Time-Series Synthesis

Excerpt

arXiv:2608.13741v1 Announce Type: cross Abstract: Synthesizing time series from natural language is emerging as the most expressive form of controllable time series generation. However, existing text-conditioned generators either take caption embeddings frozen from off-the-shelf text encoders, or adapt the encoder end-to-end, letting the denoising loss shape the embeddings only as a by-product. In either case, the conditioning representation is never deliberately matched to the signal modality,