Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
I'm looking for a small LLM to act as a Linux command assistant. I will use llama.cpp. use case: - User asks in natural language. Model outputs only t…
This may be somewhat specific to Qwen3.8 27b and the Apple M5 series, perhaps, but enough of us are running this combo that it's worth tossing out the…
Hi, everyone, I’ve just released VoxGen, a lightweight native inference engine for VoxCPM2, written in Rust and using Vulkan compute instead of Python…
Quite surprised to be beating other high quality quants. It took a lot of benchmarking to get here and we are quite pleased with these, hope they are…
I've been using Deepseek V4 Flash 0731 for a few weeks now and while I havent thrown it anything very hard, im quite happy with it. Using through anti…
Hi, I've noticed that the model often sees "garbled text" in its context. Sometimes it declare that the tools instructions are corrupted, sometimes it…
I've tried both and been having this debate with myself for the last few days, on two Asus Ascent GX10s (effectively the same as 2x DGX Spark): DeepSe…
submitted by /u/Acceptable-Cycle4645 [link] [comments]
https://preview.redd.it/6h4xc5o8l6nh1.png?width=1158&format=png&auto=webp&s=b65074b6baaa1faa2347e5259229c8ba803bcd4b Here is link to repo: https://git…
Real question, whether your are vibe coder or expert or whatever, what do you do? Read carefully each single word? Plan the next steps? Do the gym? su…
https://preview.redd.it/74bmvel9b5nh1.png?width=1602&format=png&auto=webp&s=0d0c1adaa016a486ffd97c4c466e980dc611b139 I've only recently started lookin…
This came from another thread or comment. I forgot exactly where, but the basic idea was to replace the 51B N-gram layer in Qwen 3.8 Next with a much…
I am still waiting for Llama 5, because Muse Spark will be too big for me, or just something between Glimmer and Spark https://x.com/finkd/status/2095…
Which one is for plan, whish one is for documenting, which one for writing code, which one for writing and running tests, which one [list continues] s…
I'm considering adding another A40 to my setup but my case doesn't have the space (or power) for it. It's a server chassis so there's no easy way to j…
Howdy, I was wondering if anyone knew of any turnkey low-power draw solutions to host inference with 10-20GB of VRAM? I have an 4x3090 AI GPU cluster…
I got fed up with ArtificialAnalysis 's intelligence vs. cost plots, so I made my own. This is an updated and refined follow-up to a previous post I m…
This was my first serious attempt at tuning a local LLM. I started because Qwen3.8-27B IQ3 was fast on my RTX 5080 but the coding quality disappointed…
This is getting asked from time to time, but since models changed a lot, I wanted to reask it. I'm looking for a trained model that can give short des…
Not written by AI all mistakes mine. I saw people on the subreddit saying that 3.6 works better without thinking . It made me want to know for certain…