Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
TL;DR: this paper proposes a method to fix hallucination rates to very low levels or zero by disabling neurons which contribute to hallucination. This…
Title submitted by /u/No_Algae1753 [link] [comments]
submitted by /u/No_Conversation9561 [link] [comments]
Hey guys, I tested Qwen3.8 Flash with llama.cpp from CPU-only to the full 96GB of my RTX PRO 6000. Short version: CPU-only reached 8.34 tok/s at a 2K…
Faster prompt processing on CPU. submitted by /u/jacek2023 [link] [comments]
submitted by /u/ipechman [link] [comments]
Remember to vote and comments guys ;) https://x.com/QwenDevs/status/2094389239761031591 submitted by /u/jacek2023 [link] [comments]
Sad that I only have 12gb of vram but this ik_llama is so fast submitted by /u/Needausernameplzz [link] [comments]
SlopTV: a YouTube live stream where the chat writes the programming. You type "capybara dj underwater rave", an LLM inflates it into a 400-word struct…
submitted by /u/t4a8945 [link] [comments]
Mistral is to be release a new model this summer, they still are working on it. What are your hopes? submitted by /u/always_posedge_clk [link] [commen…
According to this article Samsung has locked up the 70% of it's future ram production in contracts to companies like Microsoft, Google, and Nvidia. Ev…
I keep seeing demos of AI agents building scenes in Blender through BlenderMCP, so I tried it myself. I ran both models locally for this and picked th…
what words to use in the prompt that can (really) affect on the model behavior, and impact that hard not about what it talk about, but words that can…
So i spent considerable time trying to figure out why one of my eGPUs has degraded from x4 to x1 permanently. Yesterday while cleaning i found the cul…
I stole the reference image from a recent post on r/stablediffusion , and then asked both qwen 3.8 flash next (q4 K XL) and GLM flash (oQ4e MLX) to ch…
I am doing a lot of translation work with different languages, and German is just exceptionally well done by Qwen 3.8 27B. It is lengths ahead of GPT-…
(not written by Claude, all errors and crappy text are result of too little coffee on a Sunday morning ;) Our home server is a 2018 Thinkstation P520,…
I'm running Muse Glimmer 30B EXL3-SC 3.00bpw H4, fully resident on my 12GB VRAM GPU at 100K context with Q8\_O KV cache. It's a joy to use a dense 30B…
If you own 4xR9700 and were waiting for the model to make them shine, then I have some good news for you! It's running at 80-120 tokens/second for gen…