← Back to all articles
arXiv cs.CLSeptember 22, 2026

Low resource cross-modal alignment using HGNN to enhance speech representation

Excerpt

arXiv:2609.23191v1 Announce Type: new Abstract: Speech-text space alignment is a multimodal representation learning method consisting to map different speech and text into a shared representation space, leading to enrichment of the representation of each modality. Proposed architectures, such as SAMU-XLSR, typically follow a student/teacher framework, with the goal of fine-tuning an audio encoder to produce representations that closely match those of the text. In this way a speech representation