Latest AI/ML News

770 articles · Reddit r/LocalLLaMA

Reddit r/LocalLLaMAAug 25, 2026

submitted by /u/dd32x [link] [comments]

Reddit r/LocalLLaMAAug 25, 2026

I ran 11 uncensored variants of Gemma 4 12B that I grabbed from huggingface, sorting by downloads. 10 full abliterations plus 2 LoRA adapters which we…

Reddit r/LocalLLaMAAug 25, 2026

glm 5.3 flash While awaiting the release of the version 5.3 weights, this theory is gaining ground. OxAlpha is new GLM. submitted by /u/LegacyRemaster…

Reddit r/LocalLLaMAAug 25, 2026

submitted by /u/Decent-Hat-5807 [link] [comments]

Reddit r/LocalLLaMAAug 25, 2026

We have been working on some performance optimisations for Qwen3.8 and other models. The main new feature that we introduced is adaptive speculation f…

Reddit r/LocalLLaMAAug 25, 2026

submitted by /u/coder543 [link] [comments]

Reddit r/LocalLLaMAAug 25, 2026

"A 12-core GPU, also with two more cores than before, now includes Neural Accelerators in each core for the first time on Mac mini, resulting in up to…

Reddit r/LocalLLaMAAug 25, 2026

At $10k, you could get - 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan) - 5.7B tokens with DeepSeek V4 Pro OpenRouter - 100B tokens with DeepSeek V4 Fl…

Reddit r/LocalLLaMAAug 25, 2026

I spotted the Dual B60 48GB listed on Digitec/Galaxus. Initially it was said these wouldn't go into standard retail channels. At CHF 2500 (post tax, U…

Reddit r/LocalLLaMAAug 25, 2026

submitted by /u/RuthlessCriticismAll [link] [comments]

Reddit r/LocalLLaMAAug 25, 2026

I just can't resist submitted by /u/close_Meal6005 [link] [comments]

Reddit r/LocalLLaMAAug 25, 2026

Granite-4.2-30B is the flagship reasoning model in the Granite 4.2 family. It delivers the strongest performance across reasoning-intensive tasks by l…

Reddit r/LocalLLaMAAug 25, 2026

Prepare your disk space guys submitted by /u/jacek2023 [link] [comments]

Reddit r/LocalLLaMAAug 25, 2026

lpddr5x probably, the m7 ultra if is using ddr6 should be at 1.8 Tb/s submitted by /u/Last-Owl-8342 [link] [comments]

Reddit r/LocalLLaMAAug 25, 2026

submitted by /u/rerri [link] [comments]

Reddit r/LocalLLaMAAug 25, 2026

Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate: Ideal 4-bit quant ≈ 82 GB (58 GB main weights + 24 GB n-gram tables) Real-world quants li…

Reddit r/LocalLLaMAAug 25, 2026

submitted by /u/themixtergames [link] [comments]

Reddit r/LocalLLaMAAug 24, 2026

Share what your favorite models are right now and why . Given the nature of the beast in evaluating VLMs (untrustworthiness of benchmarks, immature to…

Reddit r/LocalLLaMAAug 22, 2026

I'm building a budget AI PC for our company's application. The specs are: MSI Z370 TOMAHAWK 64 GB RAM NZXT C1200 Gold PSU 2x RTX 3090 build was finish…

Reddit r/LocalLLaMAAug 22, 2026

If you use these cards together, you would think Q8 would be faster tps because it uses less VRAM. Well the reality is: Prompt TPS is nearly identical…