Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
Hello, I've been self-hosting LLMs on various budget hardware for a while (6x RTX 3060 12 GB, Intel Arc Pro B60 24 GB, RX 9070 XT, etc). Over the last…
Hi everyone, it's been a while since I posted so here's an update on what the Lemonade community has been up to this summer. Our overall mission is to…
I measured various Qwen3.8 27B quantizations by Unsloth on popular benchmarks: FPQA Diamond, IFBench, and Terminal-Bench-2.1. Q4_K_M is all you need.…
The news about Qwen 3.8 Flash Next is the first I'm reading about n-gram tables. I may be completely misunderstanding how they work but it seems they…
In the past during model releases there used to be tons of interesting discussions happening on this subreddit. However, the new rules of forcing ever…
I think that mall models between 10B and 40B are the real gamechanger. These models will be the ones that will make the AI buble pop and big companies…
Coval has a public TTS leaderboard with 25 APIs on it. The harness is open source, so we just ran it against our API at https://www.nineninesix.ai and…
Hi all, I'm looking for the best model for a hobby project and trying to make sense of the various data I came across. I know benchmarks do not often…
Hugging Face is exploring sale of the business valued at around $13 billion dollars. Actually I don't think we have any other repo source. Which has t…
Like wtaf? Qwen 3.8 27b is crazy. Can't wait for kimi k3 performance submitted by /u/GrokiniGPT [link] [comments]
Megathread for discussing the release of GLM-5.3-Flash. Quants Fine-Tunes & Abliterations Chat Templates Inference Server Support & Configuration Expe…
The Artist: Qwen3.8-27B-UD-Q3\ K_XL, q8_0 caches, xhigh, temp 1.0, image-min-tokens 1024, froggeric template) I was screwing around with different Qwe…
submitted by /u/coder543 [link] [comments]
I wanted to see just how capable Qwen3.8-27b is locally. I have a RTX 4090 and 96GB of RAM but the Q4 comfortably fits in the GPU with plenty of conte…
submitted by /u/BriguePalhaco [link] [comments]
we made vision mlx quants of qwen3.8 27b (9 builds from 8bit at 29.5 GB down to 3.23bpw DWQ at 11.8 GB) and compared them against other community visi…
The company on Wednesday confirmed speculation that the Ox Alpha model is a new iteration of its GLM series and said it will release the weights for i…
submitted by /u/pscoutou [link] [comments]
In the past, I have use llama.cpp, but I read that the exl3 quantization format should give better precision , so I have tried exllamav3/tabbyAPI. It…
Benchmarked qwen3.8 xhigh, medium and muse glimmer. Xhigh effort mode with qwen3.8 took almost 30hrs. (And still failed on 16 cases because of the 32K…