Latest AI/ML News

770 articles · Reddit r/LocalLLaMA

Reddit r/LocalLLaMASep 26, 2026

I've been playing around with Ling 3.0 Tiny, which is an 8 billion parameter model (MoE, 1B active). And I've had a lot of poignant thoughts as a resu…

Reddit r/LocalLLaMASep 26, 2026

For those of us who have been around for a while, we witnessed the huge demand for desktops, servers, and specialized appliances in the early 2000s. E…

Reddit r/LocalLLaMASep 26, 2026

I posted an updated GPT-OSS template a couple of months ago , which was based on Unsloth's version . It turns out that both Unsloth's version (and, th…

Reddit r/LocalLLaMASep 26, 2026

Hello r/localllama once again, it's me your kobold concedo Been a few months since I last posted here, and today I have something new I'd like to shar…

Reddit r/LocalLLaMASep 26, 2026

I wish Qwen also released dataset and method to fully train a model ourselves but it is what it is. However, I come here with my stupid question becau…

Reddit r/LocalLLaMASep 26, 2026

I am new to playing around with local ai and have a ton to learn about it but I was curious, why are they (who is they?) releasing them for free, don'…

Reddit r/LocalLLaMASep 26, 2026

Enjoy! Fucking loving it. submitted by /u/sleight42 [link] [comments]

Reddit r/LocalLLaMASep 26, 2026

These people are so Naive to think like this without any knowledge. submitted by /u/pmv143 [link] [comments]

Reddit r/LocalLLaMASep 26, 2026

Today I present a fine piece of engineering, carefully assembled inside a custom chipboard chassis: the 2400cc Inference Racer , a.k.a. my winter heat…

Reddit r/LocalLLaMASep 24, 2026

This has made a massive improvement in performance on my 7900XTX before: | model | size | params | backend | ngl | n_ubatch | fa | test | t/s | | ----…

Reddit r/LocalLLaMASep 24, 2026

The CoT of MinimMax M3.1, currently available in openrouter and opencode under the guise of "Space Bunny Alpha", has the familiar look of caveman mode…

Reddit r/LocalLLaMASep 24, 2026

Model Overview Ling-3.0-flash-VL inherits the language, reasoning, and long-context capabilities of Ling-3.0-flash, while extending them with native i…

Reddit r/LocalLLaMASep 24, 2026

I posted previously my swe-verified django 100 tasks benchmark comparing different local models and quantization. No new models for now, but a fix in…

Reddit r/LocalLLaMASep 24, 2026

Many a praise have been sung on Qwen-3.8, but here is mine. Qwen-3.8 and I had a rocky start, because it thinks so much. Watching it working is painfu…

Reddit r/LocalLLaMASep 24, 2026

Hello Everyone! Good news for anyone who can compile code on apple, you can now experience using a AMD Radeon AI Pro R9700 on a MacBook Pro or Mac min…

Reddit r/LocalLLaMASep 24, 2026

These boards cost me $115 each and I have them connected using llama.cpp with Vulkan and RPC on Bazzite. The boards have roughly 27GB of combined GPU…

Reddit r/LocalLLaMASep 24, 2026

Ok, so the Mac M5 Ultra (256GB) hit the market, but the only publicly available benchmarks material are flashy YouTube "clown influencers" videos. We…

Reddit r/LocalLLaMASep 24, 2026

Original post: https://www.reddit.com/r/LocalLLaMA/comments/1woscea/contrastive_language_models/ (sorry I felt it wasn't giving CLM the highlight it d…

Reddit r/LocalLLaMASep 23, 2026

i keep seeing people recommending, for people with Blackwell/ 5090 cards, that using nvfp4 is "a no brainer" because of the big speed benefits. but my…

Reddit r/LocalLLaMASep 24, 2026

For Qwen3.8 27B - Q8 with 262k context, or higher with YaRN. Using llamacpp or better Thanks :) submitted by /u/evillarreal86 [link] [comments]