Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
Sharing my harness for running local LLMs that I built using Qwen 3.x 27B (> 90% locally built) under my supervision - not vibe-coded. Its free, no te…
How much system RAM? How much VRAM? How much SSD space? Ideally list for q3/4 but q2 might also work since I have seen 3.8 27B perform well even on q2…
The promise has been fulfilled. submitted by /u/serige [link] [comments]
submitted by /u/ElementNumber6 [link] [comments]
UD 3.0 seems to be a massive improvement over UD 2.0 Some of us still want to run the older Qwen models but would benefit from UD 3.0 UD 2.0 vs 3.0 is…
specs hardware: M4 Max 128GB Studio inference engine: oMLX & lllama.cpp insights it still very early, so had to disable oMLX K/V caching, qwen4_exp ar…
https://www.businessinsider.com/nvidia-in-talks-to-buy-hugging-face-13-billion-dollars-2026-8 Edit: The Information is reporting that the deal is done…
Repost because reddit keeps thinking this is piracy or illegal. It is neither. A lot of people are skeptical Nvidia will keep huggingface intact now t…
submitted by /u/Normal-Phone7762 [link] [comments]
I don't know why but when watching a video about FreeToken this morning this just came to mind lol. submitted by /u/MammothUnique4147 [link] [comments…
submitted by /u/johnnyApplePRNG [link] [comments]
submitted by /u/XMasterrrr [link] [comments]
Hi r/LocalLLaMA ! We’re Apodex , the team behind Apodex 1.1 , our new model family built to scale agentic intelligence for complex work. We’re excited…
*part 2 of an earlier post: previous quant comparison with voxel island creation this time I rented three rtx pro 6000 96gb, on each one I launched a…
I'd be curious to try it locally since I use 0731 daily but still no news on the weights submitted by /u/LegacyRemaster [link] [comments]
Forgive the typos and rambling - human actually wrote this post 😂 I've had a Strix Halo board for about a year and been playing around with it for va…
Upgraded to a 5070 Ti so I could run Qwen 3.8 27B, which works perfectly, but didn't want to let the old 4070 Ti go to waste. The cards would touch if…
I was bored and handwrote a tiny 100-line bash script to let an agent search for and read articles from an offline wikipedia archive during a regular…
submitted by /u/dreamai87 [link] [comments]
Yet another vLLM fork thread here, but this time its for older INT8-centric hardware. This is a complete INT8 serving stack for Qwen3.8 27B based on v…