Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
Here’s hoping RTX 6090 also comes at MSRP of $6969 lol. With how expensive RTX 5090s and 6000 Pros have gotten it doesn’t sound too far fetched. submi…
All this mega threads and censoring posts killing LocalLLaMA vibe. And yeah, I liked more when we had 20+ posts about new model. submitted by /u/inkbe…
I was using UD-Q3_K_XL until now with more than 140000 context. Quality wise it's very good, very few erroneous tool calls. Then I saw many others her…
submitted by /u/f0urxio [link] [comments]
I’ve been working on a small project called gemma4.c. The idea is pretty simple: you can download a modern language model, compile one 700-line C file…
With HF being bought out and its future feeling a little iffy, I got to thinking about the teams that have constantly looked out for the little guys a…
finally I can download the GGUF UPDATE Q4 GGUF downloaded, I have 55 t/s on 4x3090, video in the comment submitted by /u/jacek2023 [link] [comments]
Ever since Qwen 3.8 Flash Next dropped, there's a misconception going around that N-gram tables will let people run 1T+ parameter models on a single s…
I was planning on another 5090, but then I realize... perhaps I am much better off getting an M5 Ultra Mac Studio with 256gb of ram. We are so genuine…
With this move Nvidia is not only acquiring the HuggingFace platform, but they might also effectively acquire the copyright to the llama.cpp project,…
Same as title. If we could do that can we potentially do expert caching from ram to vram so it's faster? If not the router itself, can we train a smal…
Spent yesterday getting Qwen3.8 Flash and GLM 5.3 Flash up and running on my cluster of 4 x DGX Sparks with a view to replacing DeepSeek 0731... but..…
Qwen models are both pareto frontiers in total size AND active parameters size of all open weights models so far. If this trend continues, we might se…
HuggingFace releases microduck a 10 inch open-source biped with 15 actuators and sensors (camera, speaker, LiDAR, NFC, bluetooth, wifi, ...) that you…
There is an absolutely insane amount of waste downloading entire finetunes if we can just load Loras on top of base models instead. submitted by /u/Bo…
Qwen 3.8 Next Flash (Qwen 4) engrams don't need to be in VRAM/RAM submitted by /u/jacek2023 [link] [comments]
Pollen Robotics and Hugging Face are releasing an open-source bipedal robot that comes with reinforcement learning software. It looks like it has a sp…
Hey all! I'm finally doing some cool stuff with my "thinking heater" (h/t u/-TV-Stand- ). I'm still experimenting with GLM-5.2 (in anticipation of 5.3…
tl;dr faster dense models for low VRAM people option similar to the existing --n-cpu-moe It puts user specified amount of FFN sublayers for dense mode…
Since Qwen's dropped the Qwen4Exp architecture bomb that focus on offloading parameters to n-gram instead of pure mixture of experts, I dug into this…