Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
> "Qwen 3.8 is a damn good coder, but a terrible collaborator" It modifies SO many things in my scripts for what should be a 2 line PR, I get a 100 li…
Looking at ... some of the new memory architecture. ... I hired my good friend, Seok-Hee Lee, who used to run SK Hynix. ... We are not ready to unfold…
One GPU drops off immediately when vllm is started and the other throws CUDA errors on start submitted by /u/cantgetthistowork [link] [comments]
Extended reasoning and post-training appear to be the keys used by DeepSeek, Qwen, and GLM to boost performance (leveraging higher token counts). And…
These are screenshots from the r/Singularity comment section. I'm speechless. This doesn't even have downvotes. How can someone cheer for a monopoly r…
What’s y’all’s best guess on parameter size based on these weird-ass names? submitted by /u/Porespellar [link] [comments]
One of the other posts today by user u/Howard_banister confirmed what I've been seeing from the other AI subreddits as well. Most of these other subs…
I used Claude Code to help write this pipeline - Gemma4 to turn an idea into a short story, then IndexTTS 2.5 running on an NVIDIA 3080 to use voices…
Add 4 experts into Dense model and confirmed recovering model's ability up to "general level". Hey, google. Please releae official 124B MoE model!!!!!…
Just saw this. https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp submitted by /u/Key_Solid_1696 [link] [comments]
Normally, whenever a new model dropped, I always chose the non-vision version just to save VRAM; I though that only use case was when you were the one…
I haven't had a chance to test it yet, but it looks very promising. It seems to speed up MTP for MoE models across different draft widths (especially…
submitted by /u/paf1138 [link] [comments]
A hallucinated event, Crono awakens in his modest bedroom of 2095... I am using local models since 2025 January. My daily driver is qwen 3.6 35B A3B,…
Interesting new paper from Alexia Jolicoeur-Martineau (of Tiny Recursive Model fame) and collaborators. They seem to be able to replace quadratic atte…
Hi all, I have recently been experimenting with different LLM set ups and after everyone was raving about how good Qwen 3.8 27b was, I was inspired to…
I've made a modification of llama.cpp MTP for people that want to run models like QWEN 27B on 16GB and similar setup, the focus is reducing the memory…
When using kvarn quants at low context depth tg speed is similar to llama.cpp on and equivalent qx_x quant. However, as context depth grows kvarn tank…
https://huggingface.co/Nanbeige/Nanbeige4.2-3B-DSpark 4b model for the gpu poor that I think is stronger than qwen 3.5 9b now faster! I was getting ro…
With b10726, the default --lazy-mode change keeps the 51B-parameter PLE n-gram embedding table of Qwen 3.8 Flash Next on disk: it is mmap'd and its ro…