Latest AI/ML News

770 articles · Reddit r/LocalLLaMA

Reddit r/LocalLLaMAAug 22, 2026

Im looking to set up a new local llm (probably on unsloth studio as that seemed to be doing pretty well last time I tested it). This one won't need to…

Reddit r/LocalLLaMAAug 22, 2026

I get why they exist and in almost mostly any other hardware field we can see clearly the difference and what it respects throughout, but with ai, its…

Reddit r/LocalLLaMAAug 22, 2026

I have a gigabyte ds3h v2 b450 motherboard. Currently hosting an RTX 3090 I have a spare 3070 and I wondered, can I run both? Got myself a riser cable…

Reddit r/LocalLLaMAAug 22, 2026

I love the Ornith 35B local models, 1.0 has been running my HAM radio rig for me. I have a hackRF receiver and a 5 watt quansheng portable the both ru…

Reddit r/LocalLLaMAAug 22, 2026

Hello, I have a few questions that I can't seem to find a clear answer to. Does it make sense to make your own GGUF? I noticed that when I compile lla…

Reddit r/LocalLLaMAAug 22, 2026

For example: Instead of saying: "I created this new ID" It says: "I minted this new ID" Instead of: "This alternative path is available" It says: "thi…

Reddit r/LocalLLaMAAug 22, 2026

This is a genuine community of real generally respectful adult human beings. Despite the enthusiasm all of you have for local AI, you can recognize th…

Reddit r/LocalLLaMAAug 22, 2026

I was experimenting with deepseek harness when found that even if you don't use deepseek models, you can configure the web_search tool with their api…

Reddit r/LocalLLaMAAug 22, 2026

Using llama.cpp I seem to be unable to get my to GPUs working tougether correclty, so I need help somehow. Setup: 96GB RAM, one Blackwell 5000 (48GB)…

Reddit r/LocalLLaMAAug 22, 2026

submitted by /u/johnnyApplePRNG [link] [comments]

Reddit r/LocalLLaMAAug 22, 2026

New/Old benchmark that provides a lot of answers for local LLM. I present to you a new test that I developed somewhat by accident: https://huggingface…

Reddit r/LocalLLaMAAug 22, 2026

Idea is simple-ish in abstract: instead of running normal prefill over all of a given prompt, split it into parts - generates caches for part A and pa…

Reddit r/LocalLLaMAAug 22, 2026

I was finally able to replicate tensor level allocation outside the Gemma family. https://huggingface.co/ByteOtter/Qwen3.5-4B-CADA-IQ2_XS After the Ge…

Reddit r/LocalLLaMAAug 22, 2026

No GGUFs yet on Huggingface though. submitted by /u/Ihtien [link] [comments]

Reddit r/LocalLLaMAAug 22, 2026

A new project was released yesterday and I have the opportunity to test it today. Papper: https://arxiv.org/abs/2608.16157 Github: https://github.com/…

Reddit r/LocalLLaMAAug 22, 2026

https://preview.redd.it/84zi5nsdawkh1.png?width=2368&format=png&auto=webp&s=1109e69db807b153064b1f5b61d22cf1e9fbca05 Another user posted the benchmark…

Reddit r/LocalLLaMAAug 22, 2026

Has anyone else experimented with ubatch size when running GLM-5.2 locally? I was testing the 226 GiB GLM-5.2-UD-IQ2\_XXS GGUF on 3x RTX PRO 6000 Blac…

Reddit r/LocalLLaMAAug 22, 2026

You can find the changelog and source code here: https://github.com/ggml-org/llama.cpp/releases/tag/v0.2.0 Associated pre-build is here: https://githu…