arXiv cs.CLSeptember 21, 2026
Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
Excerpt
arXiv:2511.13254v2 Announce Type: replace Abstract: Large Language Models (LLMs) have displayed remarkable capabilities across diverse domains, but their training remains resource- and time-intensive, requiring massive computational resources and careful orchestration of training procedures. Model souping-the practice of averaging weights from multiple models of the same architecture-has emerged as a promising pre- and post-training technique that can enhance performance without expensive retrai