Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
Quote: DeepSeek is training a 2T-parameter model and plans to eventually build an 8T-parameter model. https://x.com/wallstengine/status/21019828436563…
submitted by /u/Bestlife73 [link] [comments]
I’ve been working on making small models more capable at agentic coding and work, because most people in the world don’t have the sort of hardware nee…
submitted by /u/fallingdowndizzyvr [link] [comments]
submitted by /u/themixtergames [link] [comments]
https://huggingface.co/yandex/AliceAI-Foundation-80B-A3B-Base It's not a Qwen3 finetune, it's actually its own fully custom architecture. No Llama.cpp…
Hey everyone! It has been quite a while since the last SupraLabs model - but today we've something special for y'all: Supra2-IMG It's a 100M parameter…
This sub is, needless to say very niche and skewed towards the high end. There are tons of extremely high end setups here with multiple gpu's etc. Eve…
submitted by /u/Bestlife73 [link] [comments]
submitted by /u/Hyacin75 [link] [comments]
submitted by /u/johnnyApplePRNG [link] [comments]
After seeing a few posts on here about it, I finally tried exl3 3bpw and exllamav3 for running flash next - with amazing results. On 3x3090s, 128GB DD…
I have been reasonably satisfied with my single R9700 (32GB) as I can run practical quants of Qwen 3.8-27B at good speeds, as well as other similar mo…
I know cloud subscription plans are not exactly the core focus of r/LocalLLaMA , so I want to be clear about why I’m posting this here. I’ve been buil…
Another day, another Qwen Flash Next speedup submitted by /u/jacek2023 [link] [comments]
After seeing u/Nandakishor_ml’s post introducing Laya , I wanted to see how fast it could run in a standalone C++ implementation. Credit to u/Nandakis…
submitted by /u/johnnyApplePRNG [link] [comments]
Laya is an open-weight (Apache 2.0) "System 1" decision model from Convai Innovations, built by Nandakishor M as an open alternative to TypeSafe's clo…
so I was doing some stuffs in Deepseek harness with Qwen 3.8 27b and just glanced over to see the progress and saw this, make me chuckle 😄 submitted…
I’ve spent basically the last 8 hours testing different models on the exact same web-development prompt, and I finally finished. The whole point of th…