Latest AI/ML News

770 articles · Reddit r/LocalLLaMA

Reddit r/LocalLLaMASep 24, 2026

https://preview.redd.it/knblecgpeerh1.png?width=2048&format=png&auto=webp&s=9f4b2e85ab3c2a9961ead7eb27f392d0d6001f92 submitted by /u/simulated-souls […

Reddit r/LocalLLaMASep 24, 2026

Source: https://rogueon.ai/blog/rogue-deep-think-preview/tecnica Could this kinda stuff be the way for open models to prevail? Since open models are s…

Reddit r/LocalLLaMASep 23, 2026

submitted by /u/Fcking_Chuck [link] [comments]

Reddit r/LocalLLaMASep 24, 2026

I just lurk here mostly. i enjoy my 30tps on q3 3.8 27b and stop messing with it, but recently ive heard people talking about ngrams and im curious ab…

Reddit r/LocalLLaMASep 24, 2026

Just watched a video from three years ago where Matt Berman tested Bard (yes, remember Bard?) and most of the tests were like summarization or logic t…

Reddit r/LocalLLaMASep 23, 2026

Got QFN up and running on our Strix box this past weekend and have been running on it for a few days now. Big thank you/shout out to the Halogen team,…

Reddit r/LocalLLaMASep 23, 2026

Thought it's about time to share after testing for a week. You need four things most people miss: the right quant, the right model, the right branch,…

Reddit r/LocalLLaMASep 24, 2026

I keep reading about extreme ends of TPS what people consider running their models at. Some are happy to run huge quality models at merely sub 30 tps,…

Reddit r/LocalLLaMASep 23, 2026

Gemma 5, 220B A18B QAT plus ngrams please. Thank you very much! Will settle for 120B A16B plus ngrams. submitted by /u/CriticallyCarmelized [link] [co…

Reddit r/LocalLLaMASep 23, 2026

I have a lot of projects with my friends and team at work that I copy to use for my personal projects, whether it's a plugin I borrow with their conse…

Reddit r/LocalLLaMASep 23, 2026

This keeps evolving and they seem to be appearing and disappearing by the day. Is LMStudio's Bionic any good? submitted by /u/DrDisintegrator [link] […

Reddit r/LocalLLaMASep 23, 2026

I've been running DeepSeek-V4-Flash-0731 on two Radeon AI PRO R9700s (32 GB each, 192 GB system RAM) using affinity( https://codeberg.org/StillDeadcod…

Reddit r/LocalLLaMASep 23, 2026

TLDR; Base M5 Ultra 96 GB ran Q3.8 FN aggregate 3.2k PP and ~170 TG in 4 concurrency Alert: Numbers and custom server details at end are AI assisted S…

Reddit r/LocalLLaMASep 23, 2026

submitted by /u/dryadofelysium [link] [comments]

Reddit r/LocalLLaMASep 23, 2026

Hi, this is the research project I'm working on. I must say, building a GUI harness is way more complicated than I thought. The goal is something like…

Reddit r/LocalLLaMASep 23, 2026

I hear Qwen code unlocks the model better. I also think it has more power user features than open code? It’s nice open code can work with multiple mod…

Reddit r/LocalLLaMASep 23, 2026

If like me you enjoyed classics like Baldur's Gate 2, this is a small fun experiment. For anyone interested: https://www.youtube.com/live/8FhPfRKTucw?…

Reddit r/LocalLLaMASep 23, 2026

I am proud to release these quants of Qwen 3.8 27B. They beat the excellent ISTA and Unsloth quants byte-for-byte on three corpora. Both KLD and top 1…

Reddit r/LocalLLaMASep 23, 2026

The rumor about Kimi execs getting arrested finally has some legs. I believe the reality is more like under investigation for potential arrests or fin…