arXiv cs.LGOctober 7, 2026
Investigating Model Compression for Neural Machine Translation in the Biomedical Domain
Excerpt
arXiv:2610.07032v1 Announce Type: cross Abstract: Large-scale pretrained transformer models have achieved state-of-the-art performance across diverse machine translation tasks, including multilingual settings. Knowledge distillation has emerged as a sustainable approach for model compression, transferring knowledge from large teacher models to smaller, more efficient student models. Similarly, quantization, which reduces the numerical precision of model weights and activations (e.g., from 32-bit