Latest AI/ML News

770 articles · Reddit r/LocalLLaMA

Reddit r/LocalLLaMASep 13, 2026

I have a full model, it's ready to train. It's ~9b parameters. 9.4b to be more exact. That includes a 1/2/3 Engram table, Moonshot's AttnRes modeling,…

Reddit r/LocalLLaMASep 13, 2026

Using the RTX3090 on a linux machine I built for it and running Qwen 3.8 27B getting average 100t/s compared to my MacBook 20t/s I think I can finally…

Reddit r/LocalLLaMASep 13, 2026

This is an evolution on top of Raymond's KV cache streaming fork - all credits to what enabled this goes to him. The basic idea behind what he enabled…

Reddit r/LocalLLaMASep 13, 2026

So literally two days ago I discovered pi-vcc from someone's comment reply in this sub. I installed it, and we're off to the races. Sub-second compact…

Reddit r/LocalLLaMASep 13, 2026

I'm not fond of having an irreplaceable subscription on some OpenAI and Anthropic and being eventually being unable to do my job without using tools f…

Reddit r/LocalLLaMASep 13, 2026

I can get $5k for the 5090 and the Mac is $5499 before tax. The 5090 has a memory bandwidth of 1.8 TB/s while the M5 Ultra is 1.2 TB/s. Is this a sens…

Reddit r/LocalLLaMASep 13, 2026

from internlm: We introduce Intern-S2-397B , our most capable multimodal foundation model for scientific intelligence and long-horizon agents. Intern-…

Reddit r/LocalLLaMASep 14, 2026

I've been thinking of translating some light novels. Have 24GB VRAM. submitted by /u/RadianceTower [link] [comments]

Reddit r/LocalLLaMASep 13, 2026

I couldn't find any posts mentioned this windows build via the subreddit search. Based on ZLUDA, but properly compiled for Windows, might be exciting.…

Reddit r/LocalLLaMASep 13, 2026

I'll start by saying I'm not talking about the models themselves, I'm aware that I can't come close to something like Fable's intelligence locally. Ju…

Reddit r/LocalLLaMASep 13, 2026

> "make an svg of a frog playing on a chello on the back of a whale with carribean island in the back." interestingly the svg looks different in the O…

Reddit r/LocalLLaMASep 13, 2026

Not sure if anyone out there is working on this, don't see any on huggingface to try out. It might be too new at the moment, but it would be cool to s…

Reddit r/LocalLLaMASep 12, 2026

submitted by /u/pmv143 [link] [comments]

Reddit r/LocalLLaMASep 13, 2026

The entire page of the NY Times today above the fold absent one article is AI (the models are just too strong/too dangerous, must be regulated). They…

Reddit r/LocalLLaMASep 13, 2026

VLLM Benchmark: Prefill, Prompt processing - avg, 871.93 tok/s (3 hours constant running xhigh) - 10K prompt, 1000.26 tok/s (16 runs) - 90K prompt, 74…

Reddit r/LocalLLaMASep 13, 2026

Does anyone think the gurus on the DGX Spark forum are going to figure out how to magically fit DeepSeek 4.1 Flash on a 2x cluster, or is it only poss…

Reddit r/LocalLLaMASep 13, 2026

It's still a dream for many folks to run medium size(30B range) models @ Q8 with Unquantized KVCache (256K Context) on their GPUs. It would be awesome…

Reddit r/LocalLLaMASep 14, 2026

https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3 AA shipped a new benchmark last week as part of the Intelligence In…

Reddit r/LocalLLaMASep 13, 2026

The first generation of our 150M model has just been released Its performance is similar to that of GPT2-Small The benchmarks: PIQA: 62.24% Hellaswag:…

Reddit r/LocalLLaMASep 13, 2026

Lately because of the current hardware shortage, unfortunately or fortunately, we can’t just throw infinite cloud compute at our problems, but we’re f…