Latest AI/ML News

770 articles · Reddit r/LocalLLaMA

Reddit r/LocalLLaMASep 2, 2026

> "Qwen 3.8 is a damn good coder, but a terrible collaborator" It modifies SO many things in my scripts for what should be a 2 line PR, I get a 100 li…

Reddit r/LocalLLaMASep 1, 2026

Looking at ... some of the new memory architecture. ... I hired my good friend, Seok-Hee Lee, who used to run SK Hynix. ... We are not ready to unfold…

Reddit r/LocalLLaMASep 2, 2026

One GPU drops off immediately when vllm is started and the other throws CUDA errors on start submitted by /u/cantgetthistowork [link] [comments]

Reddit r/LocalLLaMASep 2, 2026

Extended reasoning and post-training appear to be the keys used by DeepSeek, Qwen, and GLM to boost performance (leveraging higher token counts). And…

Reddit r/LocalLLaMASep 1, 2026

These are screenshots from the r/Singularity comment section. I'm speechless. This doesn't even have downvotes. How can someone cheer for a monopoly r…

Reddit r/LocalLLaMASep 1, 2026

What’s y’all’s best guess on parameter size based on these weird-ass names? submitted by /u/Porespellar [link] [comments]

Reddit r/LocalLLaMASep 2, 2026

One of the other posts today by user u/Howard_banister confirmed what I've been seeing from the other AI subreddits as well. Most of these other subs…

Reddit r/LocalLLaMAAug 31, 2026

I used Claude Code to help write this pipeline - Gemma4 to turn an idea into a short story, then IndexTTS 2.5 running on an NVIDIA 3080 to use voices…

Reddit r/LocalLLaMASep 1, 2026

Add 4 experts into Dense model and confirmed recovering model's ability up to "general level". Hey, google. Please releae official 124B MoE model!!!!!…

Reddit r/LocalLLaMAAug 31, 2026

Just saw this. https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp submitted by /u/Key_Solid_1696 [link] [comments]

Reddit r/LocalLLaMAAug 31, 2026

Normally, whenever a new model dropped, I always chose the non-vision version just to save VRAM; I though that only use case was when you were the one…

Reddit r/LocalLLaMAAug 31, 2026

I haven't had a chance to test it yet, but it looks very promising. It seems to speed up MTP for MoE models across different draft widths (especially…

Reddit r/LocalLLaMAAug 31, 2026

A hallucinated event, Crono awakens in his modest bedroom of 2095... I am using local models since 2025 January. My daily driver is qwen 3.6 35B A3B,…

Reddit r/LocalLLaMAAug 31, 2026

Interesting new paper from Alexia Jolicoeur-Martineau (of Tiny Recursive Model fame) and collaborators. They seem to be able to replace quadratic atte…

Reddit r/LocalLLaMAAug 31, 2026

Hi all, I have recently been experimenting with different LLM set ups and after everyone was raving about how good Qwen 3.8 27b was, I was inspired to…

Reddit r/LocalLLaMAAug 31, 2026

I've made a modification of llama.cpp MTP for people that want to run models like QWEN 27B on 16GB and similar setup, the focus is reducing the memory…

Reddit r/LocalLLaMAAug 31, 2026

When using kvarn quants at low context depth tg speed is similar to llama.cpp on and equivalent qx_x quant. However, as context depth grows kvarn tank…

Reddit r/LocalLLaMAAug 31, 2026

https://huggingface.co/Nanbeige/Nanbeige4.2-3B-DSpark 4b model for the gpu poor that I think is stronger than qwen 3.5 9b now faster! I was getting ro…

Reddit r/LocalLLaMAAug 31, 2026

With b10726, the default --lazy-mode change keeps the 51B-parameter PLE n-gram embedding table of Qwen 3.8 Flash Next on disk: it is mmap'd and its ro…