Latest AI/ML News
2649 articles · 👥 Community Buzz
Hello Everyone! Good news for anyone who can compile code on apple, you can now experience using a AMD Radeon AI Pro R9700 on a MacBook Pro or Mac min…
These boards cost me $115 each and I have them connected using llama.cpp with Vulkan and RPC on Bazzite. The boards have roughly 27GB of combined GPU…
Ok, so the Mac M5 Ultra (256GB) hit the market, but the only publicly available benchmarks material are flashy YouTube "clown influencers" videos. We…
Original post: https://www.reddit.com/r/LocalLLaMA/comments/1woscea/contrastive_language_models/ (sorry I felt it wasn't giving CLM the highlight it d…
i keep seeing people recommending, for people with Blackwell/ 5090 cards, that using nvfp4 is "a no brainer" because of the big speed benefits. but my…
For Qwen3.8 27B - Q8 with 262k context, or higher with YaRN. Using llamacpp or better Thanks :) submitted by /u/evillarreal86 [link] [comments]
https://preview.redd.it/knblecgpeerh1.png?width=2048&format=png&auto=webp&s=9f4b2e85ab3c2a9961ead7eb27f392d0d6001f92 submitted by /u/simulated-souls […
Source: https://rogueon.ai/blog/rogue-deep-think-preview/tecnica Could this kinda stuff be the way for open models to prevail? Since open models are s…
submitted by /u/Fcking_Chuck [link] [comments]
I just lurk here mostly. i enjoy my 30tps on q3 3.8 27b and stop messing with it, but recently ive heard people talking about ngrams and im curious ab…
Just watched a video from three years ago where Matt Berman tested Bard (yes, remember Bard?) and most of the tests were like summarization or logic t…
Got QFN up and running on our Strix box this past weekend and have been running on it for a few days now. Big thank you/shout out to the Halogen team,…
Thought it's about time to share after testing for a week. You need four things most people miss: the right quant, the right model, the right branch,…
I keep reading about extreme ends of TPS what people consider running their models at. Some are happy to run huge quality models at merely sub 30 tps,…
Gemma 5, 220B A18B QAT plus ngrams please. Thank you very much! Will settle for 120B A16B plus ngrams. submitted by /u/CriticallyCarmelized [link] [co…
I have a lot of projects with my friends and team at work that I copy to use for my personal projects, whether it's a plugin I borrow with their conse…
This keeps evolving and they seem to be appearing and disappearing by the day. Is LMStudio's Bionic any good? submitted by /u/DrDisintegrator [link] […
https://github.com/PersonalJarvis/PersonalJarvis submitted by /u/InternationalGap3698 [link] [comments]
I've been running DeepSeek-V4-Flash-0731 on two Radeon AI PRO R9700s (32 GB each, 192 GB system RAM) using affinity( https://codeberg.org/StillDeadcod…
TLDR; Base M5 Ultra 96 GB ran Q3.8 FN aggregate 3.2k PP and ~170 TG in 4 concurrency Alert: Numbers and custom server details at end are AI assisted S…