Latest AI/ML News

770 articles · Reddit r/LocalLLaMA

Reddit r/LocalLLaMASep 12, 2026

So, uh... the popularity of so-called humanlike Qwen (currently on top in this sub) made me realize just how clueless the general public is about the…

Reddit r/LocalLLaMASep 12, 2026

https://huggingface.co/HermiHg/Qwen3.8-27B-DFlash2-Q2_K_S-MIX-GGUF I used this draft model with https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO…

Reddit r/LocalLLaMASep 12, 2026

Hey everyone! I'm curious to hear from people that use a combination of cloud-based frontier models and local ones for development. I'm planning to se…

Reddit r/LocalLLaMASep 12, 2026

Most model leaderboards assume a server with powerful GPUs to run models that people daily use. However, my smolbenchmark is the other column: models…

Reddit r/LocalLLaMASep 12, 2026

submitted by /u/Fcking_Chuck [link] [comments]

Reddit r/LocalLLaMASep 12, 2026

submitted by /u/Thrumpwart [link] [comments]

Reddit r/LocalLLaMASep 12, 2026

AuK-Flash: Fast 4-Step Speech Generation and Editing arXiv : https://arxiv.org/abs/2609.08936 Full Paper : https://arxiv.org/pdf/2609.08936 GitHub : h…

Reddit r/LocalLLaMASep 12, 2026

I find this new model at HF: "Built for demanding work. A 262 144-token context window, adjustable reasoning effort, tool calling, and text, image and…

Reddit r/LocalLLaMASep 12, 2026

First off, I know that GLM, Qwen, and DeepSeek absolutely dominate in terms of SOTA Open Source models, and that’s what I use in my personal projects…

Reddit r/LocalLLaMASep 12, 2026

Blog Post : Per-tensor layout maps for GGUF quantization Reddit thread : New tensor type layouts for my GGUF uploads EDIT : Model card has updated thi…

Reddit r/LocalLLaMASep 12, 2026

Applied science work, from workflow design, data pipeline, results analysis, article/reports writing and data publishing online. 5 projects I did in t…

Reddit r/LocalLLaMASep 12, 2026

Coxon, bernie and now this First https://x.com/DarioAmodei/status/2098773920774074715 Then https://x.com/elonmusk/status/2098789109980332057 Then http…

Reddit r/LocalLLaMASep 11, 2026

I've been out of the loop for some time. Is there already an official way to offload ngram to RAM or SSD in something like Unsloth Studio? Interested…

Reddit r/LocalLLaMASep 11, 2026

I’ve been building "Spomin", a router that replaces context with summaries directly in the KV cache. The goal is to keep long running sessions going w…

Reddit r/LocalLLaMASep 11, 2026

Developer's own thread: https://www.reddit.com/r/LocalLLaMA/s/adp1cGZZe9 Code: https://github.com/Inovello/llama.cpp/tree/flashnext-e06 My hardware: 2…

Reddit r/LocalLLaMASep 11, 2026

I wanted to see how my fully local home voice assistant compared to the latest GPT Live, so I tested it using the same conversation used in their "Imp…

Reddit r/LocalLLaMASep 11, 2026

A huge favor to ask llama maintainers - please implement this feature. Even with one 3090 card there will be tangible improvements in decode speed on…

Reddit r/LocalLLaMASep 11, 2026

Can you still do something with a 2050 or something like it? I mean for office work, loading embedding, reranking and chat models not at the same time…

Reddit r/LocalLLaMASep 10, 2026

Github Repo. Blog post. 💡 TL;DR (from the Github Readme) Spend less without making the agent do less useful work. SoL-Pi is a standalone extension fo…

Reddit r/LocalLLaMASep 10, 2026

Like many of you, I have seen many posts and tweets in the last weeks complaining about Artificial Analysis being "broken", "meaningless", and "bought…