Reddit r/LocalLLaMASeptember 1, 2026
Given how common RTX 3090 use is for LLMs, why don't we see more INT8 W8A8 models ?
Excerpt
Based on https://huggingface.co/hardware , the RTX 3090 is the second most used GPU by LLM enthusiasts. Because RTX 3090 has native INT8 tensors cores, it can provide better performance with INT8 W8A8. However people seems to default to FP8 or smaller quants anyway. I suppose I am missing information that explains why ? submitted by /u/TheOnlyBen2 [link] [comments]