Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
I have a full model, it's ready to train. It's ~9b parameters. 9.4b to be more exact. That includes a 1/2/3 Engram table, Moonshot's AttnRes modeling,…
Using the RTX3090 on a linux machine I built for it and running Qwen 3.8 27B getting average 100t/s compared to my MacBook 20t/s I think I can finally…
This is an evolution on top of Raymond's KV cache streaming fork - all credits to what enabled this goes to him. The basic idea behind what he enabled…
So literally two days ago I discovered pi-vcc from someone's comment reply in this sub. I installed it, and we're off to the races. Sub-second compact…
I'm not fond of having an irreplaceable subscription on some OpenAI and Anthropic and being eventually being unable to do my job without using tools f…
I can get $5k for the 5090 and the Mac is $5499 before tax. The 5090 has a memory bandwidth of 1.8 TB/s while the M5 Ultra is 1.2 TB/s. Is this a sens…
from internlm: We introduce Intern-S2-397B , our most capable multimodal foundation model for scientific intelligence and long-horizon agents. Intern-…
I've been thinking of translating some light novels. Have 24GB VRAM. submitted by /u/RadianceTower [link] [comments]
I couldn't find any posts mentioned this windows build via the subreddit search. Based on ZLUDA, but properly compiled for Windows, might be exciting.…
I'll start by saying I'm not talking about the models themselves, I'm aware that I can't come close to something like Fable's intelligence locally. Ju…
> "make an svg of a frog playing on a chello on the back of a whale with carribean island in the back." interestingly the svg looks different in the O…
Not sure if anyone out there is working on this, don't see any on huggingface to try out. It might be too new at the moment, but it would be cool to s…
submitted by /u/pmv143 [link] [comments]
The entire page of the NY Times today above the fold absent one article is AI (the models are just too strong/too dangerous, must be regulated). They…
VLLM Benchmark: Prefill, Prompt processing - avg, 871.93 tok/s (3 hours constant running xhigh) - 10K prompt, 1000.26 tok/s (16 runs) - 90K prompt, 74…
Does anyone think the gurus on the DGX Spark forum are going to figure out how to magically fit DeepSeek 4.1 Flash on a 2x cluster, or is it only poss…
It's still a dream for many folks to run medium size(30B range) models @ Q8 with Unquantized KVCache (256K Context) on their GPUs. It would be awesome…
https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3 AA shipped a new benchmark last week as part of the Intelligence In…
The first generation of our 150M model has just been released Its performance is similar to that of GPT2-Small The benchmarks: PIQA: 62.24% Hellaswag:…
Lately because of the current hardware shortage, unfortunately or fortunately, we can’t just throw infinite cloud compute at our problems, but we’re f…