Latest AI/ML News

770 articles · Reddit r/LocalLLaMA

Reddit r/LocalLLaMAAug 19, 2026

https://preview.redd.it/8i4eixqpsakh1.png?width=3637&format=png&auto=webp&s=5b9501308cb045dd44949b8dd2df90b98aa36840 Have been using local models to b…

Reddit r/LocalLLaMAAug 19, 2026

The gap between "medium" and the default "xhigh" is ridiculously huge. Medium barely thinks, xhigh... well there has already been many posts about tha…

Reddit r/LocalLLaMAAug 18, 2026

From Hugging Face on 𝕏: https://x.com/huggingface/status/2089673018737869242?s=20 submitted by /u/Nunki08 [link] [comments]

Reddit r/LocalLLaMAAug 19, 2026

I have been searching for suitable model to run on my 8GB RAM toy, NVIDIA Orin Nano Super 8GB. This little toy was priced at $249 earlier this year (n…

Reddit r/LocalLLaMAAug 19, 2026

Lucebox now, or wait for the new Framework Desktop (Ryzen AI Max+ PRO 495 / 192 GB) + PCIe x4-to-x16 adapter & Radeon AI PRO R9700? submitted by /u/Mo…

Reddit r/LocalLLaMAAug 18, 2026

Two days ago I released a hyper-optimized Qwen3.8-27B inference engine for an RTX 3090 (82 tps single request, 672 peak) - yesterday's update took tha…

Reddit r/LocalLLaMAAug 18, 2026

Here's the DFlash2 announcement , and I was pretty excited for this after trying out DSpark on llama.cpp a few days ago and being somewhat disappointe…

Reddit r/LocalLLaMAAug 18, 2026

submitted by /u/johnnyApplePRNG [link] [comments]

Reddit r/LocalLLaMAAug 18, 2026

Not even sure if I'm joking, my thinking history is about 50% "wait". submitted by /u/pixelpoet_nz [link] [comments]

Reddit r/LocalLLaMAAug 19, 2026

submitted by /u/pscoutou [link] [comments]

Reddit r/LocalLLaMAAug 18, 2026

submitted by /u/coder543 [link] [comments]

Reddit r/LocalLLaMAAug 18, 2026

submitted by /u/surreal_tournament [link] [comments]

Reddit r/LocalLLaMAAug 18, 2026

am I the only one who does this lol submitted by /u/close_Meal6005 [link] [comments]

Reddit r/LocalLLaMAAug 18, 2026

Qwen released the 2.4T Max weights and I was curious how well it can re-create COD in one prompt I ran the model on a rented B200 cluster and used rou…

Reddit r/LocalLLaMAAug 18, 2026

I managed to run the 143–144 GiB DeepSeek-V4-Flash-0731 UD-Q4_K_XL GGUF on four RTX 3060 12GB cards while keeping a 360k–376k context window. Hardware…

Reddit r/LocalLLaMAAug 18, 2026

submitted by /u/johnnyApplePRNG [link] [comments]

Reddit r/LocalLLaMAAug 18, 2026

submitted by /u/anderspitman [link] [comments]

Reddit r/LocalLLaMAAug 18, 2026

submitted by /u/johnnyApplePRNG [link] [comments]

Reddit r/LocalLLaMAAug 18, 2026

Apparently a second version of DFlash from the original authors of DFlash GGUF quants are already made available with an accompanying llama.cpp PR: ht…

Reddit r/LocalLLaMAAug 18, 2026

Who needs GPUs? submitted by /u/DeltaSqueezer [link] [comments]