← Back to all articles
Reddit r/LocalLLaMASeptember 1, 2026

Given how common RTX 3090 use is for LLMs, why don't we see more INT8 W8A8 models ?

Excerpt

Based on https://huggingface.co/hardware , the RTX 3090 is the second most used GPU by LLM enthusiasts. Because RTX 3090 has native INT8 tensors cores, it can provide better performance with INT8 W8A8. However people seems to default to FP8 or smaller quants anyway. I suppose I am missing information that explains why ? submitted by /u/TheOnlyBen2 [link] [comments]