Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
I can't seem to find a good answer to this, my Hermes agent has access to Firecrawl and some other web scrapers for content extraction, but anyone kno…
I just saw the price of RXT 6000 Pro and the Thor IGX. while the price looks similar but the IGX you will get a full setup not only the GPU. Anyone lo…
Sharp v22.1 is u/peculiar-ragdoll 's system prompt that makes Qwen answer way more tersely without losing correctness. NInfer is a hyper-tailored infe…
Context: Earlier, open-source large models like Mimo V2.5 Pro, DeepSeek V4 Pro (first version), and Kimi K2.5 used to struggle with this prompt, and Q…
It was only a matter of time... submitted by /u/Retumbo77 [link] [comments]
I finally made the move from LM Studio to vLLM thanks to this post https://www.reddit.com/r/LocalLLaMA/s/NmS9CgHvqz . I may not know what it all means…
To anyone who needs AIR… submitted by /u/Miserable-Dare5090 [link] [comments]
Hello gang, I made an implementation of DSpark PC Tree (Parent conditioned drafting tree). This is an implementation of this research paper: https://a…
Hi, I have been using opencode with openrouter for quite a while now. Having read the success stories of using Qwen3.8-27B, I thought of trying it too…
I tried this model yesterday, and it felt to me like the best one I've tried for a local model for interactive use; the responses and reasoning are ve…
My Pro subscription expired today, they killed my access at 1pm local time. I'm now using Qwen3.8-27b w/ 5090m 24gb vram and pi to do everything i was…
Have seen some people say Qwen 3.8 still overthinks even when reasoning is set to low. Which on my case has been way better compared to 3.6, eveb on a…
QwQ was genuine next-gen performance usable on local hardware, but the massive required context (it's reasoning style was akin to "if I say every poss…
I wanted to just test the unsloth 1bit quant of qwen 3.8 27b as I have just 8gb vram and ngl it gave me a good laugh submitted by /u/Ok-Health-7096 [l…
Hi fellows fully-local halos, after manually following existing guides, I decided to build an LLM API endpoint installation and optimization guide tha…
Here's a brand new Blackwell-native, prefill-optimized 4-bit quant that runs 50% faster on compatible hardware than a Q4 quant of the same memory foot…
So what’s actually the best local AI harness rn? I’ve read a TON about this already and somehow ended up more confused than when I started so I figure…
dots3-note preview is the first open-weight model in the dots3 family. It is a Mixture-of-Experts model with 280B total parameters, 16B activated para…
FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation HuggingFace : https://…
That is the second comparison and the last one. I will not be spamming again ;) Continuation from: https://www.reddit.com/r/LocalLLaMA/comments/1vu0u2…