Latest AI/ML News
2649 articles · 👥 Community Buzz
How sanoTTS works? I have vibe coded this site to show what's inside sanoTTS? Every tensor shown on the page is a real intermediate value captured fro…
Gonna apologize in advance since this isn’t directly ML related. But TLDR is I’m curious what other think will happen to the value of a PhD given the…
Also Q2 and Q3 are they really that bad as people around say? or it depends highly on the model architecture and the quantize technique? is there Big…
The Nvidia GTX-1080Ti has 11GB of VRAM and capable for running 8B to 14B local models. I added an AMD Radeon MI50 16GB VRAM. Now with 27GB VRAM I'm ab…
Meet Inco Splash, open-source inference engine, built around the model and around Apple silicon. Up to 3× the decode speed of Ollama, 2× oMLX, and alm…
I would like to stay at q4, and serve the model to various coding harnesses. None of the models or tools I have tried will run without running out of…
I am trying to learn more about long-term memory and how people are handling that nowadays. Historically, I have used a mixture of paid cloud AI servi…
I tried google and on almost each mode.people complains it’s not good enough? is there any good tts right now that support English and can express emo…
I improved the TPS of Gemma 4 31B. Improving TPS and performing optimisations requires understanding of the model architecture, and I had to fork VLLM…
Play it right in your browser, no download. Some of the GLB’s are messy, but overall I’ve enjoyed playing with it, and the last iteration made was the…
This post is a follow up to this other post where I let Qwen 3.8 run for 63 hours autonomously to try to solve the Riemann hypothesis: https://www.red…
*8bit, vllm, 4x dgx* Prompt: Go download and use Remotion and create a cool 60-second motion graphics video with it. Impress me totally. The motion gr…
I needed a relatively simple but acceptable level of AI for working on one project. I didn't have any heavy requests, I just needed to give the AI acc…
TL;DR: I run a mismatched Tesla V100-PCIE pair—one 16 GB card and one 32 GB card, 48 GB total—in a Proxmox/LXC-based local-inference lab. The practica…
I did a small test of the new hyped quantisation of Qwen3.8 vs the biggest quant which fits into my limited 16GB VRAM with decent context. The results…
I was somewhat disappointed with the performance of Qwen 3.8 Flash Next on my single RTX 5090 using llama.cpp. One issue is that llama.cpp still has n…