← Back to all articles
Reddit r/LocalLLaMAAugust 26, 2026

Gemma4 31B vs Qwen3.8 27B - why the huge difference in benchmarks?

Excerpt

Hi all, I'm looking for the best model for a hobby project and trying to make sense of the various data I came across. I know benchmarks do not often translate to the real world, especially to your particular use case (whatever it may be). But this is truly baffling: AA says Qwen 3.8 27B is better by miles: https://artificialanalysis.ai/models/comparisons/qwen3-8-27b-vs-gemma-4-31b?intelligence-comparison=intelligence-vs-end-to-end-response-time While Arena says Gemma 4 31B is almost 20 places a