Latest AI/ML News

770 articles · Reddit r/LocalLLaMA

Reddit r/LocalLLaMASep 2, 2026

I'm looking for a small LLM to act as a Linux command assistant. I will use llama.cpp. use case: - User asks in natural language. Model outputs only t…

Reddit r/LocalLLaMASep 2, 2026

This may be somewhat specific to Qwen3.8 27b and the Apple M5 series, perhaps, but enough of us are running this combo that it's worth tossing out the…

Reddit r/LocalLLaMASep 2, 2026

Hi, everyone, I’ve just released VoxGen, a lightweight native inference engine for VoxCPM2, written in Rust and using Vulkan compute instead of Python…

Reddit r/LocalLLaMASep 2, 2026

Quite surprised to be beating other high quality quants. It took a lot of benchmarking to get here and we are quite pleased with these, hope they are…

Reddit r/LocalLLaMASep 3, 2026

I've been using Deepseek V4 Flash 0731 for a few weeks now and while I havent thrown it anything very hard, im quite happy with it. Using through anti…

Reddit r/LocalLLaMASep 2, 2026

Hi, I've noticed that the model often sees "garbled text" in its context. Sometimes it declare that the tools instructions are corrupted, sometimes it…

Reddit r/LocalLLaMASep 3, 2026

I've tried both and been having this debate with myself for the last few days, on two Asus Ascent GX10s (effectively the same as 2x DGX Spark): DeepSe…

Reddit r/LocalLLaMASep 3, 2026

submitted by /u/Acceptable-Cycle4645 [link] [comments]

Reddit r/LocalLLaMASep 2, 2026

https://preview.redd.it/6h4xc5o8l6nh1.png?width=1158&format=png&auto=webp&s=b65074b6baaa1faa2347e5259229c8ba803bcd4b Here is link to repo: https://git…

Reddit r/LocalLLaMASep 2, 2026

Real question, whether your are vibe coder or expert or whatever, what do you do? Read carefully each single word? Plan the next steps? Do the gym? su…

Reddit r/LocalLLaMASep 2, 2026

https://preview.redd.it/74bmvel9b5nh1.png?width=1602&format=png&auto=webp&s=0d0c1adaa016a486ffd97c4c466e980dc611b139 I've only recently started lookin…

Reddit r/LocalLLaMASep 2, 2026

This came from another thread or comment. I forgot exactly where, but the basic idea was to replace the 51B N-gram layer in Qwen 3.8 Next with a much…

Reddit r/LocalLLaMASep 2, 2026

I am still waiting for Llama 5, because Muse Spark will be too big for me, or just something between Glimmer and Spark https://x.com/finkd/status/2095…

Reddit r/LocalLLaMASep 2, 2026

Which one is for plan, whish one is for documenting, which one for writing code, which one for writing and running tests, which one [list continues] s…

Reddit r/LocalLLaMASep 2, 2026

I'm considering adding another A40 to my setup but my case doesn't have the space (or power) for it. It's a server chassis so there's no easy way to j…

Reddit r/LocalLLaMASep 2, 2026

Howdy, I was wondering if anyone knew of any turnkey low-power draw solutions to host inference with 10-20GB of VRAM? I have an 4x3090 AI GPU cluster…

Reddit r/LocalLLaMASep 2, 2026

I got fed up with ArtificialAnalysis 's intelligence vs. cost plots, so I made my own. This is an updated and refined follow-up to a previous post I m…

Reddit r/LocalLLaMASep 2, 2026

This was my first serious attempt at tuning a local LLM. I started because Qwen3.8-27B IQ3 was fast on my RTX 5080 but the coding quality disappointed…

Reddit r/LocalLLaMASep 2, 2026

This is getting asked from time to time, but since models changed a lot, I wanted to reask it. I'm looking for a trained model that can give short des…

Reddit r/LocalLLaMASep 2, 2026

Not written by AI all mistakes mine. I saw people on the subreddit saying that 3.6 works better without thinking . It made me want to know for certain…