← Back to all articles
arXiv cs.LGAugust 18, 2026

What Makes a Good Layer? Assessing the Layer-Wise Intrinsic Properties of Music Foundation Models

Excerpt

arXiv:2608.14819v1 Announce Type: cross Abstract: Music foundation models are commonly used as frozen audio feature extractors, yet selecting which layer to extract from remains largely heuristic. Current practice defaults to fixed depths or multi-layer fusion, with limited understanding of why certain layers transfer better across downstream tasks or how representation quality varies with depth and pre-training paradigm. We conduct a systematic layer-wise analysis of 12 music foundation models