← Back to all articles
arXiv cs.LGOctober 7, 2026

SSR: Sparse Segment Reduction for Ternary GEMM Acceleration

Excerpt

arXiv:2610.08403v1 Announce Type: new Abstract: Large Language Models (LLMs) require substantial computational resources, limiting their deployment on resource-constrained hardware. Ternary LLMs mitigate these demands through weight quantization via ternary values, achieving significant compression often with 50-90% sparsity. However, existing approaches have limitations: methods optimized for ternary weights, such as BitNet, redundant segment reduction (RSR), and its improved version RSR++, do