Latest AI/ML News

2649 articles · 👥 Community Buzz

Reddit r/LocalLLaMASep 28, 2026

Basically the idea is take your favorite model, for example qwen3.8-27b or say dsv4vision. Strip everything out that is not needed by that model so th…

Reddit r/LocalLLaMASep 27, 2026

I run my own LLM while also having a Openai subscription. Also tried DeepSeek (latest flash now). I run Qwen 3.8 flash Next at an amazing speed on my…

Reddit r/LocalLLaMASep 27, 2026

There is always at least 1+GB of VRAM not usable not matter how I set the --tensor-split (-ts) param. I tiny shift toward one side will move the weigh…

Reddit r/LocalLLaMASep 27, 2026

I built this, so the disclosure goes up front. It's a free, public MCP server plus a REST API with Canadian privacy law data. The MCP endpoint is at h…

Reddit r/LocalLLaMASep 28, 2026

Someone prompted different LLMs to generate CAD code for a bridge under fixed constraints (2-foot span, under 500g filament, 18-hour print limit), pri…

Reddit r/LocalLLaMASep 28, 2026

I think I need a 3-slot for my two cards. but holy fuck these things are pricey. submitted by /u/starkruzr [link] [comments]

Reddit r/LocalLLaMASep 27, 2026

Early this year when I was first looking at building up my inference capability you could get the 16GB Tesla P100s for between $60 and $80. Asked clau…

Reddit r/LocalLLaMASep 27, 2026

Just made this post for those who missed it : https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b UkisAI released their updated Qwen 27B (tuned for to…

Reddit r/LocalLLaMASep 27, 2026

I am using the Deepseek Harness, which has webfetch plugins by default. While it can help browse the internet, I myself have to give it specific URLs…

Reddit r/LocalLLaMASep 28, 2026

Apple Silicon Macs on macOS 26+ come with a small LLM built in. No download, no API key, and nothing leaves your Mac. Why I built it I was making a to…

Reddit r/LocalLLaMASep 28, 2026

I've liked how Muse-Glimmer worked, so I wanted to see if I could produce a smaller "kid" out of it. Ornith's sharp decisions on when to think and whi…

Reddit r/LocalLLaMASep 27, 2026

Finally got 3.8-Flash-Next running on my M4Pro 48GB Mac with https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF Dense 3.8-27B is just…

Reddit r/LocalLLaMASep 27, 2026

We built an inference engine for MoE models that don't fit in VRAM + RAM. Most of the model stays on the SSD, and experts are read as tokens need them…

Reddit r/LocalLLaMASep 27, 2026

submitted by /u/Automatic-Arm8153 [link] [comments]

Reddit r/LocalLLaMASep 27, 2026

https://naive.ai/en/research/ Built for coding and AI R&D 1M context Hybrid SWA/DSA submitted by /u/nullmove [link] [comments]

Reddit r/LocalLLaMASep 27, 2026

Saw the hundreds of tweets where people just keep asking Opus 5.5 for motion graphic videos. Decided to ask qwen to look at them and make its own. Qui…

Reddit r/LocalLLaMASep 27, 2026

EDIT: About that, I tried it out, and it's garbage so far. I did some basic tests through AIHubMix (do not use that platform btw, it's trash), and my…

Reddit r/LocalLLaMASep 28, 2026

Using an asrock 12 unit case running one board as the main with the rest of them headless. About 71GB of vram exposed. So far 40 tok/s is with 30k con…

Reddit r/LocalLLaMASep 27, 2026

Meta came out with a banger paper https://arxiv.org/pdf/2606.00206 , but it did not look at various quantizations supported in llama.cpp. So I did a r…

Reddit r/LocalLLaMASep 27, 2026

Been doing a bunch of vibe coding lately. Had my agents host a private WoW server for me, then built out a web browser client so you can play without…

← PreviousPage 17 of 133Next →