Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
https://preview.redd.it/knblecgpeerh1.png?width=2048&format=png&auto=webp&s=9f4b2e85ab3c2a9961ead7eb27f392d0d6001f92 submitted by /u/simulated-souls […
Source: https://rogueon.ai/blog/rogue-deep-think-preview/tecnica Could this kinda stuff be the way for open models to prevail? Since open models are s…
submitted by /u/Fcking_Chuck [link] [comments]
I just lurk here mostly. i enjoy my 30tps on q3 3.8 27b and stop messing with it, but recently ive heard people talking about ngrams and im curious ab…
Just watched a video from three years ago where Matt Berman tested Bard (yes, remember Bard?) and most of the tests were like summarization or logic t…
Got QFN up and running on our Strix box this past weekend and have been running on it for a few days now. Big thank you/shout out to the Halogen team,…
Thought it's about time to share after testing for a week. You need four things most people miss: the right quant, the right model, the right branch,…
I keep reading about extreme ends of TPS what people consider running their models at. Some are happy to run huge quality models at merely sub 30 tps,…
Gemma 5, 220B A18B QAT plus ngrams please. Thank you very much! Will settle for 120B A16B plus ngrams. submitted by /u/CriticallyCarmelized [link] [co…
I have a lot of projects with my friends and team at work that I copy to use for my personal projects, whether it's a plugin I borrow with their conse…
This keeps evolving and they seem to be appearing and disappearing by the day. Is LMStudio's Bionic any good? submitted by /u/DrDisintegrator [link] […
https://github.com/PersonalJarvis/PersonalJarvis submitted by /u/InternationalGap3698 [link] [comments]
I've been running DeepSeek-V4-Flash-0731 on two Radeon AI PRO R9700s (32 GB each, 192 GB system RAM) using affinity( https://codeberg.org/StillDeadcod…
TLDR; Base M5 Ultra 96 GB ran Q3.8 FN aggregate 3.2k PP and ~170 TG in 4 concurrency Alert: Numbers and custom server details at end are AI assisted S…
submitted by /u/dryadofelysium [link] [comments]
Hi, this is the research project I'm working on. I must say, building a GUI harness is way more complicated than I thought. The goal is something like…
I hear Qwen code unlocks the model better. I also think it has more power user features than open code? It’s nice open code can work with multiple mod…
If like me you enjoyed classics like Baldur's Gate 2, this is a small fun experiment. For anyone interested: https://www.youtube.com/live/8FhPfRKTucw?…
I am proud to release these quants of Qwen 3.8 27B. They beat the excellent ISTA and Unsloth quants byte-for-byte on three corpora. Both KLD and top 1…
The rumor about Kimi execs getting arrested finally has some legs. I believe the reality is more like under investigation for potential arrests or fin…