Latest AI/ML News

770 articles · Reddit r/LocalLLaMA

Reddit r/LocalLLaMASep 9, 2026

I run these models on Kaggle notebook, so not all TTS models, such as the ones that use conda env, are compatible (Or I just haven't found a way for t…

Reddit r/LocalLLaMASep 9, 2026

Cosmos3 INT4 T2I + I2V on Apple Silicon — code, weights and a Grok comparison GitHub - https://github.com/gtrg55/cosmos3-quant-mlx-cuda HF weights - h…

Reddit r/LocalLLaMASep 9, 2026

mentria.ai is a browser inference engine I've been building solo, from scratch in WebGPU/WGSL. This week it crossed a milestone I had been chasing for…

Reddit r/LocalLLaMASep 9, 2026

Model available here: Qwen3.8-27B-Uncensored-Genesis-V1-MTP-GGUF This model is a practical realisation of things described in this paper, but adapted…

Reddit r/LocalLLaMASep 9, 2026

Just a reminder for those out there itching to get into local llms - don't let FOMO or "gear acquisition syndrom" take over. No matter the hobby, it's…

Reddit r/LocalLLaMASep 9, 2026

I have been a dwarfstar fan for awhile and I really liked glm 5.3 flash but needed it to be materially faster to feel good using it. In the screenshot…

Reddit r/LocalLLaMASep 9, 2026

Amd announced this: https://www.amd.com/en/products/workstations/amd-threadripper-halo-station.html submitted by /u/Apprehensive_Bar6609 [link] [comme…

Reddit r/LocalLLaMASep 9, 2026

A few posts tagged with "new model" present models that are finetunes. My opinion : I'd rather have the "new model" tag reserved for new "major" relea…

Reddit r/LocalLLaMASep 9, 2026

Who is the marketing genius at LM Studio that decided that going ALL IN on pushing their new Bionic Agent product meant they are going to make it a gi…

Reddit r/LocalLLaMASep 9, 2026

submitted by /u/Few_Painter_5588 [link] [comments]

Reddit r/LocalLLaMASep 7, 2026

I'm currently running qwen 3.8 27b q8 and super happy with it. I have a DGX spark so could theoretically run at bf16 quantization. I know the measurab…

Reddit r/LocalLLaMASep 7, 2026

I wanted a quick calories counter for myself, using LLMs to evaluate the calories from pictures of meals + descriptions. I needed to pick a model so I…

Reddit r/LocalLLaMASep 7, 2026

I finally pulled the plug and got a second 5080 and 64GB DDR5 RAM to add to my existing 64. Now I have dual RTX5080s and 128Gb DDR5 RAM on a consumer…

Reddit r/LocalLLaMASep 7, 2026

Since the Qwen3.5 0.8B model is an interesting one for small specialized fine tunes, I was curious how fast it can run on CPUs. Why CPUs? Mainly becau…

Reddit r/LocalLLaMASep 7, 2026

My XFX Radeon RX 7900 GRE 16GB Vram GPU struggles with models over 20B size. I added my Radeon RX 480 8GB Vram GPU to the system and ran a few benchma…

Reddit r/LocalLLaMASep 7, 2026

Long time user of 3.6 27b, switched over to Next Flash since it's a logical step up even from 3.8 27b. It's soooo verbose, i'm talking 13 minutes of t…

Reddit r/LocalLLaMASep 7, 2026

Hi, I use hermes as a harness, and I am pleased with it, but sometimes Hermes's context size is a bit too much for my system, so I wanted to delegate…

Reddit r/LocalLLaMASep 7, 2026

https://preview.redd.it/3p6234jzk2oh1.png?width=900&format=png&auto=webp&s=b9eab3351d6c6dbd4d5b3fd677a6c40b57f18167 https://huggingface.co/tencent/EVI…

Reddit r/LocalLLaMASep 7, 2026

I got 2x 20GB RTX 3080s + 128GB of DDR4 2666hz RAM (only 4 of 6 channels populated) + a Xeon 6148 I've always been a llama.cpp person and I've been ru…

Reddit r/LocalLLaMASep 7, 2026

I moght get shit on for posting this but, I feel like i don't see this being talked enough and it feels like such a waste of a good piece of software.…