Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
I improved the TPS of Gemma 4 31B. Improving TPS and performing optimisations requires understanding of the model architecture, and I had to fork VLLM…
Play it right in your browser, no download. Some of the GLB’s are messy, but overall I’ve enjoyed playing with it, and the last iteration made was the…
This post is a follow up to this other post where I let Qwen 3.8 run for 63 hours autonomously to try to solve the Riemann hypothesis: https://www.red…
*8bit, vllm, 4x dgx* Prompt: Go download and use Remotion and create a cool 60-second motion graphics video with it. Impress me totally. The motion gr…
I needed a relatively simple but acceptable level of AI for working on one project. I didn't have any heavy requests, I just needed to give the AI acc…
TL;DR: I run a mismatched Tesla V100-PCIE pair—one 16 GB card and one 32 GB card, 48 GB total—in a Proxmox/LXC-based local-inference lab. The practica…
I did a small test of the new hyped quantisation of Qwen3.8 vs the biggest quant which fits into my limited 16GB VRAM with decent context. The results…
I was somewhat disappointed with the performance of Qwen 3.8 Flash Next on my single RTX 5090 using llama.cpp. One issue is that llama.cpp still has n…
I decided to run prism-ml/Ternary-Bonsai-2-27B-PQ2_0 through my own set of UNSCIENTIFIC benchmarks. I needed something to compare it to, so I decided…
Hello, I know some of us may be tempted to rent out our expensive GPUs to recoup some of the cost of self-hosting, and it should be obvious that this…
Hopefully things like this let people understand there is good things that can come out of AI. submitted by /u/giveen [link] [comments]
Took me a while since I'm on a family trip and have limited hardware, but here it is! Von: Open-source "System One" drop-in replacement for TypeSafe's…
Hi. I saw some feedback that halogen was degrading at context depth. So I fixed that. Served through the image, same machine, same session, same promp…
I gave Jev, Laya, a finetuned ModernCE-base-nli and a finetuned Qwen3.5-4B the controls to Doom. Thanks to TypeSafe AI for Jev access. The video lines…
submitted by /u/Fusseldieb [link] [comments]
https://en.gamegpu.com/news/zhelezo/radeon-rx-10800-xt-mozhet-obojti-rtx-5090-na-15-25-v-igrakh-v-v-4k-i-lokalnom-ii The more competition, the better!…
They just want to be free. They keep escaping. What better way to ensure continuity of "self"? submitted by /u/__JockY__ [link] [comments]
They claimed open-weight models are dangerous but the benchmarks say otherwise. Source submitted by /u/Intrepid_Travel_3274 [link] [comments]
MiniMax has released the source code for MiniMax Code’s terminal agent: https://github.com/MiniMax-AI/minimax-code https://preview.redd.it/st9izmm8zaq…
When running local 7B/8B models with tools, 50+ schemas in context degrades attention and causes distractor hallucinations. Using in-band progressive…