Latest AI/ML News

770 articles · Reddit r/LocalLLaMA

Reddit r/LocalLLaMAAug 26, 2026

Hello, I've been self-hosting LLMs on various budget hardware for a while (6x RTX 3060 12 GB, Intel Arc Pro B60 24 GB, RX 9070 XT, etc). Over the last…

Reddit r/LocalLLaMAAug 26, 2026

Hi everyone, it's been a while since I posted so here's an update on what the Lemonade community has been up to this summer. Our overall mission is to…

Reddit r/LocalLLaMAAug 26, 2026

I measured various Qwen3.8 27B quantizations by Unsloth on popular benchmarks: FPQA Diamond, IFBench, and Terminal-Bench-2.1. Q4_K_M is all you need.…

Reddit r/LocalLLaMAAug 26, 2026

The news about Qwen 3.8 Flash Next is the first I'm reading about n-gram tables. I may be completely misunderstanding how they work but it seems they…

Reddit r/LocalLLaMAAug 26, 2026

In the past during model releases there used to be tons of interesting discussions happening on this subreddit. However, the new rules of forcing ever…

Reddit r/LocalLLaMAAug 26, 2026

I think that mall models between 10B and 40B are the real gamechanger. These models will be the ones that will make the AI buble pop and big companies…

Reddit r/LocalLLaMAAug 26, 2026

Coval has a public TTS leaderboard with 25 APIs on it. The harness is open source, so we just ran it against our API at https://www.nineninesix.ai and…

Reddit r/LocalLLaMAAug 26, 2026

Hi all, I'm looking for the best model for a hobby project and trying to make sense of the various data I came across. I know benchmarks do not often…

Reddit r/LocalLLaMAAug 26, 2026

Hugging Face is exploring sale of the business valued at around $13 billion dollars. Actually I don't think we have any other repo source. Which has t…

Reddit r/LocalLLaMAAug 26, 2026

Like wtaf? Qwen 3.8 27b is crazy. Can't wait for kimi k3 performance submitted by /u/GrokiniGPT [link] [comments]

Reddit r/LocalLLaMAAug 26, 2026

Megathread for discussing the release of GLM-5.3-Flash. Quants Fine-Tunes & Abliterations Chat Templates Inference Server Support & Configuration Expe…

Reddit r/LocalLLaMAAug 26, 2026

The Artist: Qwen3.8-27B-UD-Q3\ K_XL, q8_0 caches, xhigh, temp 1.0, image-min-tokens 1024, froggeric template) I was screwing around with different Qwe…

Reddit r/LocalLLaMAAug 26, 2026

submitted by /u/coder543 [link] [comments]

Reddit r/LocalLLaMAAug 26, 2026

I wanted to see just how capable Qwen3.8-27b is locally. I have a RTX 4090 and 96GB of RAM but the Q4 comfortably fits in the GPU with plenty of conte…

Reddit r/LocalLLaMAAug 26, 2026

submitted by /u/BriguePalhaco [link] [comments]

Reddit r/LocalLLaMAAug 26, 2026

we made vision mlx quants of qwen3.8 27b (9 builds from 8bit at 29.5 GB down to 3.23bpw DWQ at 11.8 GB) and compared them against other community visi…

Reddit r/LocalLLaMAAug 26, 2026

The company on Wednesday confirmed speculation that the Ox Alpha model is a new iteration of its GLM series and said it will release the weights for i…

Reddit r/LocalLLaMAAug 26, 2026

submitted by /u/pscoutou [link] [comments]

Reddit r/LocalLLaMAAug 26, 2026

In the past, I have use llama.cpp, but I read that the exl3 quantization format should give better precision , so I have tried exllamav3/tabbyAPI. It…

Reddit r/LocalLLaMAAug 26, 2026

Benchmarked qwen3.8 xhigh, medium and muse glimmer. Xhigh effort mode with qwen3.8 took almost 30hrs. (And still failed on 16 cases because of the 32K…