Latest AI/ML News

770 articles · Reddit r/LocalLLaMA

Reddit r/LocalLLaMASep 4, 2026

uh so like i gave a model like 20 senses so like yeah https://huggingface.co/heterodoxin/qwen3-8b-supermultimodal submitted by /u/AccountAntique9327 […

Reddit r/LocalLLaMASep 4, 2026

This is a follow-up to my post from yesterday (17 -> 25-29 t/s with the expert cache PR). Same box: 2x RTX 3090 on PCIe 3.0, dual Xeon E5-2696 v4, 188…

Reddit r/LocalLLaMASep 3, 2026

I want to share a short paper just published exploring a simple but surprisingly effective optimization for sparse MoE reasoning models. The idea: Ins…

Reddit r/LocalLLaMASep 3, 2026

I have been trying to squeeze TTS stack down far enough to run in a $3 chip which has 512kb of SRAM without NPU. While trying to get to that milestone…

Reddit r/LocalLLaMASep 3, 2026

submitted by /u/hedonihilistic [link] [comments]

Reddit r/LocalLLaMASep 3, 2026

submitted by /u/jd_3d [link] [comments]

Reddit r/LocalLLaMASep 3, 2026

It started as a fork of llama-swap , but I have been building it out for myself since then as a convenient tool for all my local AI needs, and by now…

Reddit r/LocalLLaMASep 3, 2026

TimesFM-3 is the third generation of Google Research's zero-shot forecasting model, and the main change from 2.5 is that it handles multivariate input…

Reddit r/LocalLLaMASep 3, 2026

I would be really curious about this especially on unified memory devices. submitted by /u/giveen [link] [comments]

Reddit r/LocalLLaMASep 3, 2026

Defined as AI exceeding human cognitive abilities. 20 years in prison. Plenty of local models already fall under that big of an umbrella in some capac…

Reddit r/LocalLLaMASep 3, 2026

I start with some informations gathered thorough endless posts reading on this sub and online: Inference and hardware optimization projects https://dw…

Reddit r/LocalLLaMASep 3, 2026

Hello, So the original open deep research project from Langfuse is archived, are there any other alternatives to use with self hosted models + kiwix a…

Reddit r/LocalLLaMASep 3, 2026

Hello all, How do you deal with preventing future errors of you agents? I have made a skill which fires everytime it does something wrong. So far it i…

Reddit r/LocalLLaMASep 3, 2026

https://preview.redd.it/tiiyv2u76bnh1.png?width=3980&format=png&auto=webp&s=0501c976744e19655666b16301768878db3eda88 x.com/jukan05/status/209535308230…

Reddit r/LocalLLaMASep 3, 2026

A potato can create a very cool RPG in 24 minutes Laptop I3 8gb RAM 0gb VRAM, Windows 11 llama-server.exe --host 0.0.0.0 --port 8080 -m models\qwen3.6…

Reddit r/LocalLLaMASep 3, 2026

submitted by /u/DustNearby2848 [link] [comments]

Reddit r/LocalLLaMASep 3, 2026

We have been working on an open-source, model-neutral agent harness for general purpose agents called TrueForge, and wanted to understand how much the…

Reddit r/LocalLLaMASep 3, 2026

Ling-3.0-flash-Fin is the first finance-enhanced model in the Ant Ling family. Developed by Ant Group with leading financial institutions and domain e…

Reddit r/LocalLLaMASep 3, 2026

I had to swing down to my local Microcenter yesterday and while I was browsing around the store I noticed something odd... Inventory. They must have h…

Reddit r/LocalLLaMASep 3, 2026

For a few days I've been working on creating a custom local-only harness for some work related research using Codex / GPT 5.6 Sol and the model feels…