← Back to all articles
arXiv cs.CLSeptember 21, 2026

Scaling Forced Alignment to End-User Devices

Excerpt

arXiv:2609.21145v1 Announce Type: new Abstract: The Viterbi algorithm has been previously used to perform forced alignment of audio to text to mine training data from online resources. However, many existing implementations have quadratic time and space complexity, scaling poorly to long input sequences. We propose two optimizations to address this issue. First, we apply the Hirschberg algorithm to perform the alignment in place using linear memory. Second, we model the alignment between speech