Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
submitted by /u/dd32x [link] [comments]
I ran 11 uncensored variants of Gemma 4 12B that I grabbed from huggingface, sorting by downloads. 10 full abliterations plus 2 LoRA adapters which we…
glm 5.3 flash While awaiting the release of the version 5.3 weights, this theory is gaining ground. OxAlpha is new GLM. submitted by /u/LegacyRemaster…
submitted by /u/Decent-Hat-5807 [link] [comments]
We have been working on some performance optimisations for Qwen3.8 and other models. The main new feature that we introduced is adaptive speculation f…
submitted by /u/coder543 [link] [comments]
"A 12-core GPU, also with two more cores than before, now includes Neural Accelerators in each core for the first time on Mac mini, resulting in up to…
At $10k, you could get - 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan) - 5.7B tokens with DeepSeek V4 Pro OpenRouter - 100B tokens with DeepSeek V4 Fl…
I spotted the Dual B60 48GB listed on Digitec/Galaxus. Initially it was said these wouldn't go into standard retail channels. At CHF 2500 (post tax, U…
submitted by /u/RuthlessCriticismAll [link] [comments]
I just can't resist submitted by /u/close_Meal6005 [link] [comments]
Granite-4.2-30B is the flagship reasoning model in the Granite 4.2 family. It delivers the strongest performance across reasoning-intensive tasks by l…
Prepare your disk space guys submitted by /u/jacek2023 [link] [comments]
lpddr5x probably, the m7 ultra if is using ddr6 should be at 1.8 Tb/s submitted by /u/Last-Owl-8342 [link] [comments]
submitted by /u/rerri [link] [comments]
Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate: Ideal 4-bit quant ≈ 82 GB (58 GB main weights + 24 GB n-gram tables) Real-world quants li…
submitted by /u/themixtergames [link] [comments]
Share what your favorite models are right now and why . Given the nature of the beast in evaluating VLMs (untrustworthiness of benchmarks, immature to…
I'm building a budget AI PC for our company's application. The specs are: MSI Z370 TOMAHAWK 64 GB RAM NZXT C1200 Gold PSU 2x RTX 3090 build was finish…
If you use these cards together, you would think Q8 would be faster tps because it uses less VRAM. Well the reality is: Prompt TPS is nearly identical…