arXiv cs.LGOctober 1, 2026
Training-Free Affinity Fusion of Neural and Embedding-Based Speaker Diarization
Excerpt
arXiv:2609.39162v1 Announce Type: cross Abstract: Speaker diarization systems based on speaker embeddings and neural diarization exploit complementary forms of speaker information, but their intermediate representations are not directly compatible. We introduce Training-Free Affinity Fusion (TFAF), which integrates the speaker structure inferred by a neural diarizer into an embedding-based diarization system. The neural speaker partition is used to condition local speaker representations, from w