Latest AI/ML News

770 articles · Reddit r/LocalLLaMA

Reddit r/LocalLLaMAAug 19, 2026

Howdy, I'm back again - running my favorite benchmark (it's still unsaturated for the time being so might as well!) previous runs a b https://github.c…

Reddit r/LocalLLaMAAug 19, 2026

None has undergone post-training, giving researchers flexible starting points for continued pre-training, fine-tuning, and further research. Two key h…

Reddit r/LocalLLaMAAug 20, 2026

Personally, I gave up on 3.8 and went back to 3.6. I tried base, FP8, and few NVFP4 versions, and various parameter tunes for temperature etc., as wel…

Reddit r/LocalLLaMAAug 19, 2026

Four Tesla V100s from 2017 matched my RTX 5090 on single-request Qwen 3.8 decode. Repo: https://github.com/dnv2003/v100-skinny https://i.redd.it/5ws2a…

Reddit r/LocalLLaMAAug 19, 2026

looks like GGUF files were just updated submitted by /u/jacek2023 [link] [comments]

Reddit r/LocalLLaMAAug 19, 2026

Would love to know who's letting a 9B just go ham locally, haha But in all seriousness, how many of you are keeping to manual or manual-ish dev workfl…

Reddit r/LocalLLaMAAug 19, 2026

Aloha! 🌺Introducing Ornith-1.5, a family of open-source LLMs spanning 9B Dense, 35B MoE, and 397B MoE, trained with self-improving strategies. It ach…

Reddit r/LocalLLaMAAug 19, 2026

Intermediate tokens, called "thinking" or "reasoning" actually are nothing like it. Humans do step-by-step reasoning leading to the conclusion. LLMs u…

Reddit r/LocalLLaMAAug 20, 2026

They released both GGUFs & custom llama.cpp fork today. 35B MOE in 7GB size which's good for Mobile & Edge devices(Also low memory systems). Up to 120…

Reddit r/LocalLLaMAAug 19, 2026

Anyone tried them yet? https://huggingface.co/ornith-ai/Ornith-1.5-9B https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B https://huggingface.co/ornit…

Reddit r/LocalLLaMAAug 20, 2026

tldr: went from 16~ t/s to 27~ t/s generation. got my usable context up from 220k to the full 262k without sacrificing anything. prefill also increase…

Reddit r/LocalLLaMAAug 19, 2026

llama.cpp pr #27342 adds dflash2, so i rented an rtx 6000 and ran the same four prompts through four decoding setups on qwen3.8 27B median results ove…

Reddit r/LocalLLaMAAug 19, 2026

Edit: Title says 134 tps, it's actually 138 -- keep in mind my 3090 is power limited to 250w. Three days ago I released a hyper-optimized Qwen3.8-27B…

Reddit r/LocalLLaMAAug 20, 2026

I'm testing it since release, now with UD 3.0 in my AMD R9700 with ROCm, I always read everywhere that F16 and q8_0 for KV cache are essentially the s…

Reddit r/LocalLLaMAAug 19, 2026

I've been working on a depth pruning approach and decided to try it out on the new Qwen3.8-27B model. I managed to get the model down to about 22.7B p…

Reddit r/LocalLLaMAAug 20, 2026

This is gonna seem crazy off-topic, but I saw Spider-Man the other day and couldn’t help but notice how well executed E.V. is as an agentic system. co…

Reddit r/LocalLLaMAAug 20, 2026

Like many of you I've spent the last few days throwing Qwen3.8-27B against all of my usual use-cases and personal tasks/harnesses and workflows. It's…

Reddit r/LocalLLaMAAug 19, 2026

Hey everyone! We’re releasing new Qwen3.8-27B GGUFs with 10% higher accuracy for the same size. This uses a new version of Dynamic v3.0 Unsloth Dynami…

Reddit r/LocalLLaMAAug 20, 2026

IMO the most interesting graph in AI right now. Orange = frontier. Blue = what you can run on a 32GB RAM laptop. That means a few things: The models t…

Reddit r/LocalLLaMAAug 20, 2026

Off a single prompt, given my credentials and the name of my university, qwen3.8-27b was able to successfully pull my class schedule from the kinda sh…