Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
I run these models on Kaggle notebook, so not all TTS models, such as the ones that use conda env, are compatible (Or I just haven't found a way for t…
Cosmos3 INT4 T2I + I2V on Apple Silicon — code, weights and a Grok comparison GitHub - https://github.com/gtrg55/cosmos3-quant-mlx-cuda HF weights - h…
mentria.ai is a browser inference engine I've been building solo, from scratch in WebGPU/WGSL. This week it crossed a milestone I had been chasing for…
Model available here: Qwen3.8-27B-Uncensored-Genesis-V1-MTP-GGUF This model is a practical realisation of things described in this paper, but adapted…
Just a reminder for those out there itching to get into local llms - don't let FOMO or "gear acquisition syndrom" take over. No matter the hobby, it's…
I have been a dwarfstar fan for awhile and I really liked glm 5.3 flash but needed it to be materially faster to feel good using it. In the screenshot…
Amd announced this: https://www.amd.com/en/products/workstations/amd-threadripper-halo-station.html submitted by /u/Apprehensive_Bar6609 [link] [comme…
A few posts tagged with "new model" present models that are finetunes. My opinion : I'd rather have the "new model" tag reserved for new "major" relea…
Who is the marketing genius at LM Studio that decided that going ALL IN on pushing their new Bionic Agent product meant they are going to make it a gi…
submitted by /u/Few_Painter_5588 [link] [comments]
I'm currently running qwen 3.8 27b q8 and super happy with it. I have a DGX spark so could theoretically run at bf16 quantization. I know the measurab…
I wanted a quick calories counter for myself, using LLMs to evaluate the calories from pictures of meals + descriptions. I needed to pick a model so I…
I finally pulled the plug and got a second 5080 and 64GB DDR5 RAM to add to my existing 64. Now I have dual RTX5080s and 128Gb DDR5 RAM on a consumer…
Since the Qwen3.5 0.8B model is an interesting one for small specialized fine tunes, I was curious how fast it can run on CPUs. Why CPUs? Mainly becau…
My XFX Radeon RX 7900 GRE 16GB Vram GPU struggles with models over 20B size. I added my Radeon RX 480 8GB Vram GPU to the system and ran a few benchma…
Long time user of 3.6 27b, switched over to Next Flash since it's a logical step up even from 3.8 27b. It's soooo verbose, i'm talking 13 minutes of t…
Hi, I use hermes as a harness, and I am pleased with it, but sometimes Hermes's context size is a bit too much for my system, so I wanted to delegate…
https://preview.redd.it/3p6234jzk2oh1.png?width=900&format=png&auto=webp&s=b9eab3351d6c6dbd4d5b3fd677a6c40b57f18167 https://huggingface.co/tencent/EVI…
I got 2x 20GB RTX 3080s + 128GB of DDR4 2666hz RAM (only 4 of 6 channels populated) + a Xeon 6148 I've always been a llama.cpp person and I've been ru…
I moght get shit on for posting this but, I feel like i don't see this being talked enough and it feels like such a waste of a good piece of software.…