Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
So, uh... the popularity of so-called humanlike Qwen (currently on top in this sub) made me realize just how clueless the general public is about the…
https://huggingface.co/HermiHg/Qwen3.8-27B-DFlash2-Q2_K_S-MIX-GGUF I used this draft model with https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO…
Hey everyone! I'm curious to hear from people that use a combination of cloud-based frontier models and local ones for development. I'm planning to se…
Most model leaderboards assume a server with powerful GPUs to run models that people daily use. However, my smolbenchmark is the other column: models…
submitted by /u/Fcking_Chuck [link] [comments]
submitted by /u/Thrumpwart [link] [comments]
AuK-Flash: Fast 4-Step Speech Generation and Editing arXiv : https://arxiv.org/abs/2609.08936 Full Paper : https://arxiv.org/pdf/2609.08936 GitHub : h…
I find this new model at HF: "Built for demanding work. A 262 144-token context window, adjustable reasoning effort, tool calling, and text, image and…
First off, I know that GLM, Qwen, and DeepSeek absolutely dominate in terms of SOTA Open Source models, and that’s what I use in my personal projects…
Blog Post : Per-tensor layout maps for GGUF quantization Reddit thread : New tensor type layouts for my GGUF uploads EDIT : Model card has updated thi…
Applied science work, from workflow design, data pipeline, results analysis, article/reports writing and data publishing online. 5 projects I did in t…
Coxon, bernie and now this First https://x.com/DarioAmodei/status/2098773920774074715 Then https://x.com/elonmusk/status/2098789109980332057 Then http…
I've been out of the loop for some time. Is there already an official way to offload ngram to RAM or SSD in something like Unsloth Studio? Interested…
I’ve been building "Spomin", a router that replaces context with summaries directly in the KV cache. The goal is to keep long running sessions going w…
Developer's own thread: https://www.reddit.com/r/LocalLLaMA/s/adp1cGZZe9 Code: https://github.com/Inovello/llama.cpp/tree/flashnext-e06 My hardware: 2…
I wanted to see how my fully local home voice assistant compared to the latest GPT Live, so I tested it using the same conversation used in their "Imp…
A huge favor to ask llama maintainers - please implement this feature. Even with one 3090 card there will be tangible improvements in decode speed on…
Can you still do something with a 2050 or something like it? I mean for office work, loading embedding, reranking and chat models not at the same time…
Github Repo. Blog post. 💡 TL;DR (from the Github Readme) Spend less without making the agent do less useful work. SoL-Pi is a standalone extension fo…
Like many of you, I have seen many posts and tweets in the last weeks complaining about Artificial Analysis being "broken", "meaningless", and "bought…