Latest AI/ML News

770 articles · Reddit r/LocalLLaMA

Reddit r/LocalLLaMASep 19, 2026

I improved the TPS of Gemma 4 31B. Improving TPS and performing optimisations requires understanding of the model architecture, and I had to fork VLLM…

Reddit r/LocalLLaMASep 19, 2026

Play it right in your browser, no download. Some of the GLB’s are messy, but overall I’ve enjoyed playing with it, and the last iteration made was the…

Reddit r/LocalLLaMASep 19, 2026

This post is a follow up to this other post where I let Qwen 3.8 run for 63 hours autonomously to try to solve the Riemann hypothesis: https://www.red…

Reddit r/LocalLLaMASep 19, 2026

*8bit, vllm, 4x dgx* Prompt: Go download and use Remotion and create a cool 60-second motion graphics video with it. Impress me totally. The motion gr…

Reddit r/LocalLLaMASep 19, 2026

I needed a relatively simple but acceptable level of AI for working on one project. I didn't have any heavy requests, I just needed to give the AI acc…

Reddit r/LocalLLaMASep 19, 2026

TL;DR: I run a mismatched Tesla V100-PCIE pair—one 16 GB card and one 32 GB card, 48 GB total—in a Proxmox/LXC-based local-inference lab. The practica…

Reddit r/LocalLLaMASep 19, 2026

I did a small test of the new hyped quantisation of Qwen3.8 vs the biggest quant which fits into my limited 16GB VRAM with decent context. The results…

Reddit r/LocalLLaMASep 19, 2026

I was somewhat disappointed with the performance of Qwen 3.8 Flash Next on my single RTX 5090 using llama.cpp. One issue is that llama.cpp still has n…

Reddit r/LocalLLaMASep 19, 2026

I decided to run prism-ml/Ternary-Bonsai-2-27B-PQ2_0 through my own set of UNSCIENTIFIC benchmarks. I needed something to compare it to, so I decided…

Reddit r/LocalLLaMASep 19, 2026

Hello, I know some of us may be tempted to rent out our expensive GPUs to recoup some of the cost of self-hosting, and it should be obvious that this…

Reddit r/LocalLLaMASep 19, 2026

Hopefully things like this let people understand there is good things that can come out of AI. submitted by /u/giveen [link] [comments]

Reddit r/LocalLLaMASep 19, 2026

Took me a while since I'm on a family trip and have limited hardware, but here it is! Von: Open-source "System One" drop-in replacement for TypeSafe's…

Reddit r/LocalLLaMASep 19, 2026

Hi. I saw some feedback that halogen was degrading at context depth. So I fixed that. Served through the image, same machine, same session, same promp…

Reddit r/LocalLLaMASep 20, 2026

I gave Jev, Laya, a finetuned ModernCE-base-nli and a finetuned Qwen3.5-4B the controls to Doom. Thanks to TypeSafe AI for Jev access. The video lines…

Reddit r/LocalLLaMASep 19, 2026

https://en.gamegpu.com/news/zhelezo/radeon-rx-10800-xt-mozhet-obojti-rtx-5090-na-15-25-v-igrakh-v-v-4k-i-lokalnom-ii The more competition, the better!…

Reddit r/LocalLLaMASep 19, 2026

They just want to be free. They keep escaping. What better way to ensure continuity of "self"? submitted by /u/__JockY__ [link] [comments]

Reddit r/LocalLLaMASep 19, 2026

They claimed open-weight models are dangerous but the benchmarks say otherwise. Source submitted by /u/Intrepid_Travel_3274 [link] [comments]

Reddit r/LocalLLaMASep 18, 2026

MiniMax has released the source code for MiniMax Code’s terminal agent: https://github.com/MiniMax-AI/minimax-code https://preview.redd.it/st9izmm8zaq…

Reddit r/LocalLLaMASep 18, 2026

When running local 7B/8B models with tools, 50+ schemas in context degrades attention and causes distractor hallucinations. Using in-band progressive…