Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
Is it worth switching to from Ornith 1.5 9B? Or is there another similarly sized model that's beating both of them? (I head K2 is good but the KV cach…
I wanted to share a quick update and performance video running the full Moonshot AI Kimi K3 (moonshotai/Kimi-K3) model across my 16x GB10 cluster.…
I am seeing this JEV everywhere since yesterday in Localllama and it is passing past my head on what it is? So like what is it? Some new LLM? Or is it…
So this jev thingy is getting kind of big... tbh it seems overhyped by a large margin, but here we are. Not that its bad, just feels like we usually i…
TL;DR: Local agent loop, ~21 days, one RTX 3090. Task was pretty much "build a CUDA inference engine for optimized for yourself on this GPU arch." Got…
https://x.com/QwenDevs/status/2101917379785838660 submitted by /u/Bestlife73 [link] [comments]
You can simply run any GGUF with llama.cpp with n_predict=1 and n_probs=10, disable reasoning, and prompt it such as "If the following email is spam,…
To test what it can do. Qwen3.8-Flash-Next Intel Autoround W4A16 running locally on 4xV620 ~2k prefill and 70ts decode.. Were running around 3 hours.…
submitted by /u/fallingdowndizzyvr [link] [comments]
So my brother and I both use LLMs for coding. I've started using a local GLM 5.3 Flash instance - q4 qat. My brother uses GPT-6-Astra as his daily dri…
Sorry for the pretentious name, I know, I know.. It just contains all the pieces I would like to see a AGI model to have, and I can't stand the tempta…
Just find it interesting, since if Taalas tried to get into the consumer market, it could be interesting. Kind of like how we have game discs on CDs.…
Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨 A unified model for bo…
ZCode is now open source , and the reported security issues have been addressed. Source code: https://github.com/zai-org/ZCode The repo includes its d…
Also Q2 and Q3 are they really that bad as people around say? or it depends highly on the model architecture and the quantize technique? is there Big…
The Nvidia GTX-1080Ti has 11GB of VRAM and capable for running 8B to 14B local models. I added an AMD Radeon MI50 16GB VRAM. Now with 27GB VRAM I'm ab…
Meet Inco Splash, open-source inference engine, built around the model and around Apple silicon. Up to 3× the decode speed of Ollama, 2× oMLX, and alm…
I would like to stay at q4, and serve the model to various coding harnesses. None of the models or tools I have tried will run without running out of…
I am trying to learn more about long-term memory and how people are handling that nowadays. Historically, I have used a mixture of paid cloud AI servi…
I tried google and on almost each mode.people complains it’s not good enough? is there any good tts right now that support English and can express emo…