Latest AI/ML News

770 articles · Reddit r/LocalLLaMA

Reddit r/LocalLLaMAAug 27, 2026

Here’s hoping RTX 6090 also comes at MSRP of $6969 lol. With how expensive RTX 5090s and 6000 Pros have gotten it doesn’t sound too far fetched. submi…

Reddit r/LocalLLaMAAug 27, 2026

All this mega threads and censoring posts killing LocalLLaMA vibe. And yeah, I liked more when we had 20+ posts about new model. submitted by /u/inkbe…

Reddit r/LocalLLaMAAug 27, 2026

I was using UD-Q3_K_XL until now with more than 140000 context. Quality wise it's very good, very few erroneous tool calls. Then I saw many others her…

Reddit r/LocalLLaMAAug 27, 2026

I’ve been working on a small project called gemma4.c. The idea is pretty simple: you can download a modern language model, compile one 700-line C file…

Reddit r/LocalLLaMAAug 27, 2026

With HF being bought out and its future feeling a little iffy, I got to thinking about the teams that have constantly looked out for the little guys a…

Reddit r/LocalLLaMAAug 27, 2026

finally I can download the GGUF UPDATE Q4 GGUF downloaded, I have 55 t/s on 4x3090, video in the comment submitted by /u/jacek2023 [link] [comments]

Reddit r/LocalLLaMAAug 27, 2026

Ever since Qwen 3.8 Flash Next dropped, there's a misconception going around that N-gram tables will let people run 1T+ parameter models on a single s…

Reddit r/LocalLLaMAAug 27, 2026

I was planning on another 5090, but then I realize... perhaps I am much better off getting an M5 Ultra Mac Studio with 256gb of ram. We are so genuine…

Reddit r/LocalLLaMAAug 27, 2026

With this move Nvidia is not only acquiring the HuggingFace platform, but they might also effectively acquire the copyright to the llama.cpp project,…

Reddit r/LocalLLaMAAug 27, 2026

Same as title. If we could do that can we potentially do expert caching from ram to vram so it's faster? If not the router itself, can we train a smal…

Reddit r/LocalLLaMAAug 27, 2026

Spent yesterday getting Qwen3.8 Flash and GLM 5.3 Flash up and running on my cluster of 4 x DGX Sparks with a view to replacing DeepSeek 0731... but..…

Reddit r/LocalLLaMAAug 27, 2026

Qwen models are both pareto frontiers in total size AND active parameters size of all open weights models so far. If this trend continues, we might se…

Reddit r/LocalLLaMAAug 27, 2026

HuggingFace releases microduck a 10 inch open-source biped with 15 actuators and sensors (camera, speaker, LiDAR, NFC, bluetooth, wifi, ...) that you…

Reddit r/LocalLLaMAAug 27, 2026

There is an absolutely insane amount of waste downloading entire finetunes if we can just load Loras on top of base models instead. submitted by /u/Bo…

Reddit r/LocalLLaMAAug 27, 2026

Qwen 3.8 Next Flash (Qwen 4) engrams don't need to be in VRAM/RAM submitted by /u/jacek2023 [link] [comments]

Reddit r/LocalLLaMAAug 27, 2026

Pollen Robotics and Hugging Face are releasing an open-source bipedal robot that comes with reinforcement learning software. It looks like it has a sp…

Reddit r/LocalLLaMAAug 27, 2026

Hey all! I'm finally doing some cool stuff with my "thinking heater" (h/t u/-TV-Stand- ). I'm still experimenting with GLM-5.2 (in anticipation of 5.3…

Reddit r/LocalLLaMAAug 27, 2026

tl;dr faster dense models for low VRAM people option similar to the existing --n-cpu-moe It puts user specified amount of FFN sublayers for dense mode…

Reddit r/LocalLLaMAAug 27, 2026

Since Qwen's dropped the Qwen4Exp architecture bomb that focus on offloading parameters to n-gram instead of pure mixture of experts, I dug into this…