Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
Using an asrock 12 unit case running one board as the main with the rest of them headless. About 71GB of vram exposed. So far 40 tok/s is with 30k con…
Meta came out with a banger paper https://arxiv.org/pdf/2606.00206 , but it did not look at various quantizations supported in llama.cpp. So I did a r…
Been doing a bunch of vibe coding lately. Had my agents host a private WoW server for me, then built out a web browser client so you can play without…
It had such a long run. It was my first introduction to modern language models. I remember getting slightly excited over it. And now it lives purely i…
submitted by /u/chocolateUI [link] [comments]
Which one is better for difficult tasks like web scrapping, coding, using tools? Looking for any benchmarks because i couldn't actually find one after…
Every rent-vs-buy thread I read has confident people on both sides, but not many actually show the numbers. So I finally ran the numbers for our own d…
faster CPU prompt processing: "TL;DR: 3-7x faster CPU mul_mat using VNNI with IMO minimal complexity" submitted by /u/jacek2023 [link] [comments]
My local prompts take twenty minutes to three hours as of now. Curious for others what are you doing while it’s working? submitted by /u/xiraov [link]…
I maintain TensorSharp , an open-source inference engine. It can now run Qwen-Image 2.1 locally for text-to-image generation and image editing, with s…
https://huggingface.co/internlm/Intern-Decision-0.8B Update: https://huggingface.co/internlm/Intern-Decision-2B Intern-Decision-4B Demo | Model Weight…
Mica v0.1 4B playing a real Minecraft 1.20.4 server. Video attached. How it works - Each step the bot's live game state (inventory, nearby blocks, ent…
The conceited little fuckers love to inundate you with unnecessary details, noisy caveats, what's 'load bearing' and what's not, waste your time with…
I've been experimenting with ways to run MoE models that don't fit comfortably in RAM, and I ended up making Overspill, a disk tier for FreeToken . Th…
Maybe you'll like it? I hope I get to use my self-promotion credit a tiny little bit here after being in the community so long haha. I was the top of…
Hi, My local AI server consists of 96GB Ddr5 and one rtx 3090. I got plenty of stuff running, like krea2, qwen image 2.1, minimax h3, ltx 2.5, qwen 3.…
Sometimes it’s easy to forget that this sub and others like it are probably the extreme minority when it comes to this hobby. Most people, I would thi…
On my M5 Pro 64GB I can comfortably work in an agentic setup with the Qwen3.8 27B model in good quality (Unsloth UD-Q4_K_XL) at a decent speed of 50 t…
For instance whenever a model comes out, what’s the best engine to run it, the best harness and absolute minimum you need to get same or near same re…
About a month ago I was having FOMO and was going to spend coin I don't really have on new graphics cards. Instead of doing that though I decided to s…