Latest AI/ML News

770 articles · Reddit r/LocalLLaMA

Reddit r/LocalLLaMAAug 22, 2026

What models and configs are we using? Please share here On windows, I am using this copium pared down model https://huggingface.co/Bucoid/Qwen3.8-27B-…

Reddit r/LocalLLaMAAug 21, 2026

So usually I avoid Q3 quants because I have had bad experiences with it, models were usually too degraded, so the smallest I normally do is Q4, since…

Reddit r/LocalLLaMAAug 21, 2026

https://github.com/deepseek-ai/deepseek-harness/releases/tag/dsh-v0.1.1-rc.1 The DeepSeek adapter adds the multimodal visual understanding model DeepS…

Reddit r/LocalLLaMAAug 21, 2026

submitted by /u/Xhehab_ [link] [comments]

Reddit r/LocalLLaMAAug 21, 2026

I used LM Studio Bionic with Qwen 3.8 27B Q3_K_S with 57k context. It took a staggering 63 hours to finish coding. After the first prompt "Create a be…

Reddit r/LocalLLaMAAug 21, 2026

Even the low preset is better than Qwen 3.7 plus or Qwen3.6-27B reasoning submitted by /u/Tall_Abrocoma_3533 [link] [comments]

Reddit r/LocalLLaMAAug 21, 2026

Artificial Analysis just benchmarked them and the scores are crazy good, proving the earlier success wasn't only enabled by overthinking. submitted by…

Reddit r/LocalLLaMAAug 21, 2026

A quick feedback after a really major test: nearly 20 hours of non-stop goal-oriented work with Qwen3.8-27B Q6, running across an RTX 3090 and an RTX…

Reddit r/LocalLLaMAAug 20, 2026

IQ quants are particularly slow on CPU at large batch sizes (what you'd see for imatrix and perplexity) Benchmark numbers I ran PPL against master and…

Reddit r/LocalLLaMAAug 20, 2026

AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB GPU card — without quantization, distillatio…

Reddit r/LocalLLaMAAug 20, 2026

I see https://github.com/linuxid10t/llama.cpp/tree/feature/g9v3-support but yeah... Considering that some results place it above Qwen 3.6 27B (the top…

Reddit r/LocalLLaMAAug 20, 2026

I'm introducing Aurora-80K, a small language model with exactly 80 thousand parameters. It uses a factorized 4,096-token vocabulary despite having onl…

Reddit r/LocalLLaMAAug 20, 2026

Hey everyone, Posted TinySearch here a few versions ago and got a bunch of useful feedback, so figured I'd post an update because the thing has change…

Reddit r/LocalLLaMAAug 20, 2026

Component Validated configuration Motherboard ASRock Rack SPC621D8U-2T/OVH CPU Xeon Gold 6330 (Get gold/platinum if interested in Optane Pmem gimmicks…

Reddit r/LocalLLaMAAug 20, 2026

The pelican on a bicycle is sooo outdated, so I came up with a new, improved version. Qwen3.8-27b medium (UD-Q4_K_XL) vs. Sol 5.6 high vs. Qwen3.6-35B…

Reddit r/LocalLLaMAAug 20, 2026

From the screenshots: Hy4 is now live, labeled "Expert-Level Model" + "Use Tools to Solve Problems" Hy3 is tagged with "New Upgrade," positioned as a…

Reddit r/LocalLLaMAAug 20, 2026

I pre-trained a 1.02-billion-parameter on Kimi K3 replica trained on 5.00 billion decontaminated tokens for $250. This model has 1.02 billion paramete…

Reddit r/LocalLLaMAAug 20, 2026

Not everyone has the disposable income to build a small data center, so making this post for the underdogs as I was very surprised by the performance/…

Reddit r/LocalLLaMAAug 19, 2026

https://x.com/liquidai/status/2090078070929760295 https://huggingface.co/LiquidAI/LFM2.5-2.6B-GGUF submitted by /u/jacek2023 [link] [comments]