Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
I've been playing around with Ling 3.0 Tiny, which is an 8 billion parameter model (MoE, 1B active). And I've had a lot of poignant thoughts as a resu…
For those of us who have been around for a while, we witnessed the huge demand for desktops, servers, and specialized appliances in the early 2000s. E…
I posted an updated GPT-OSS template a couple of months ago , which was based on Unsloth's version . It turns out that both Unsloth's version (and, th…
Hello r/localllama once again, it's me your kobold concedo Been a few months since I last posted here, and today I have something new I'd like to shar…
I wish Qwen also released dataset and method to fully train a model ourselves but it is what it is. However, I come here with my stupid question becau…
I am new to playing around with local ai and have a ton to learn about it but I was curious, why are they (who is they?) releasing them for free, don'…
Enjoy! Fucking loving it. submitted by /u/sleight42 [link] [comments]
These people are so Naive to think like this without any knowledge. submitted by /u/pmv143 [link] [comments]
Today I present a fine piece of engineering, carefully assembled inside a custom chipboard chassis: the 2400cc Inference Racer , a.k.a. my winter heat…
This has made a massive improvement in performance on my 7900XTX before: | model | size | params | backend | ngl | n_ubatch | fa | test | t/s | | ----…
The CoT of MinimMax M3.1, currently available in openrouter and opencode under the guise of "Space Bunny Alpha", has the familiar look of caveman mode…
Model Overview Ling-3.0-flash-VL inherits the language, reasoning, and long-context capabilities of Ling-3.0-flash, while extending them with native i…
I posted previously my swe-verified django 100 tasks benchmark comparing different local models and quantization. No new models for now, but a fix in…
Many a praise have been sung on Qwen-3.8, but here is mine. Qwen-3.8 and I had a rocky start, because it thinks so much. Watching it working is painfu…
Hello Everyone! Good news for anyone who can compile code on apple, you can now experience using a AMD Radeon AI Pro R9700 on a MacBook Pro or Mac min…
These boards cost me $115 each and I have them connected using llama.cpp with Vulkan and RPC on Bazzite. The boards have roughly 27GB of combined GPU…
Ok, so the Mac M5 Ultra (256GB) hit the market, but the only publicly available benchmarks material are flashy YouTube "clown influencers" videos. We…
Original post: https://www.reddit.com/r/LocalLLaMA/comments/1woscea/contrastive_language_models/ (sorry I felt it wasn't giving CLM the highlight it d…
i keep seeing people recommending, for people with Blackwell/ 5090 cards, that using nvfp4 is "a no brainer" because of the big speed benefits. but my…
For Qwen3.8 27B - Q8 with 262k context, or higher with YaRN. Using llamacpp or better Thanks :) submitted by /u/evillarreal86 [link] [comments]