Latest AI/ML News

770 articles · Reddit r/LocalLLaMA

Reddit r/LocalLLaMAAug 26, 2026

https://preview.redd.it/kbsqh6f7molh1.png?width=730&format=png&auto=webp&s=068dbea9a50be634a369d54d8b27b781d020fab3 My experience with Qwen 3.8 for ag…

Reddit r/LocalLLaMAAug 26, 2026

https://x.com/romanchernin/status/2092488160680751437?s=20 - Multimodal (Vision) - 1M Tokens Context Window - DeepSWE ~63% Edit: He deleted it, screen…

Reddit r/LocalLLaMAAug 26, 2026

Megathread for discussing the (impending) release of Qwen 3.8 Flash Next. Quants Fine-Tunes & Abliterations Chat Templates Inference Server Support &…

Reddit r/LocalLLaMAAug 26, 2026

So here's the thing, almost everyone use NVIDIA to run their LLMs, we also do the same, a lot of people we've met use like RTX PRO 6000 or even H100,…

Reddit r/LocalLLaMAAug 26, 2026

submitted by /u/RedditUsr2 [link] [comments]

Reddit r/LocalLLaMAAug 26, 2026

We're releasing a fully quantized NVFP4 version of Qwen3.8-27B. The checkpoint was trained using quantization-aware distillation (QAD) with QUASAR, ou…

Reddit r/LocalLLaMAAug 26, 2026

https://preview.redd.it/i8rjx0ar5mlh1.png?width=2160&format=png&auto=webp&s=c2588bc7b2519ea71b176ca73faf566dfc585496 I wanted to see whether a heavily…

Reddit r/LocalLLaMAAug 25, 2026

I’ve been going back and forth on this for a week and I can’t settle it, so I’m hoping someone here has hands-on numbers. The two configs (German pric…

Reddit r/LocalLLaMAAug 25, 2026

Portable rig for Qwen3.8-27B-BF16 200K+ token prompts. My work Panasonic Toughbook + the T1 + power brick + headphones all fit in my lunchbox. Need th…

Reddit r/LocalLLaMAAug 25, 2026

Generated on a 3090 Qwen 3.8 27b Q4 Thinking high. submitted by /u/Both_Opportunity5327 [link] [comments]

Reddit r/LocalLLaMAAug 25, 2026

submitted by /u/indicava [link] [comments]

Reddit r/LocalLLaMAAug 25, 2026

With hopes of a Qwen3.8-35B-A3B release now mostly dashed, many people including myself are looking at fine-tunes and other variants of Qwen3.6-35B-A3…

Reddit r/LocalLLaMAAug 25, 2026

mrburns_excellent.gif submitted by /u/funding__secured [link] [comments]

Reddit r/LocalLLaMAAug 25, 2026

My best bet is the embedding-offloaded linear where the 51b n-grams track semantics and context, like 3.8 with its loss of real world knowledge bolted…

Reddit r/LocalLLaMAAug 25, 2026

WeMM-Embedding-9B is a universal multimodal embedding model built on Qwen3.5. It accepts text, images, videos, visual documents, and interleaved multi…

Reddit r/LocalLLaMAAug 25, 2026

Does the model have fill in middle support? I would love to have a smart model doing good code suggestions (auto complete). No full slop mode, just a…

Reddit r/LocalLLaMAAug 25, 2026

.. didn't llama.cpp aka ggml get acquired by Huggingface not too long ago? How would this sale affect llama.cpp and ggml? What are possible risks, and…

Reddit r/LocalLLaMAAug 24, 2026

Credit to Twitter Post submitted by /u/Rymssss [link] [comments]

Reddit r/LocalLLaMAAug 24, 2026

talking to any white collar employee submitted by /u/edge_compute_user [link] [comments]

Reddit r/LocalLLaMAAug 25, 2026

Hello everyone. I hope is all well with you all. I been around this sub for a few months and been quietly reading and I have seen how many of you are…