Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
https://preview.redd.it/8i4eixqpsakh1.png?width=3637&format=png&auto=webp&s=5b9501308cb045dd44949b8dd2df90b98aa36840 Have been using local models to b…
The gap between "medium" and the default "xhigh" is ridiculously huge. Medium barely thinks, xhigh... well there has already been many posts about tha…
From Hugging Face on 𝕏: https://x.com/huggingface/status/2089673018737869242?s=20 submitted by /u/Nunki08 [link] [comments]
I have been searching for suitable model to run on my 8GB RAM toy, NVIDIA Orin Nano Super 8GB. This little toy was priced at $249 earlier this year (n…
Lucebox now, or wait for the new Framework Desktop (Ryzen AI Max+ PRO 495 / 192 GB) + PCIe x4-to-x16 adapter & Radeon AI PRO R9700? submitted by /u/Mo…
Two days ago I released a hyper-optimized Qwen3.8-27B inference engine for an RTX 3090 (82 tps single request, 672 peak) - yesterday's update took tha…
Here's the DFlash2 announcement , and I was pretty excited for this after trying out DSpark on llama.cpp a few days ago and being somewhat disappointe…
submitted by /u/johnnyApplePRNG [link] [comments]
Not even sure if I'm joking, my thinking history is about 50% "wait". submitted by /u/pixelpoet_nz [link] [comments]
submitted by /u/pscoutou [link] [comments]
submitted by /u/coder543 [link] [comments]
submitted by /u/surreal_tournament [link] [comments]
am I the only one who does this lol submitted by /u/close_Meal6005 [link] [comments]
Qwen released the 2.4T Max weights and I was curious how well it can re-create COD in one prompt I ran the model on a rented B200 cluster and used rou…
I managed to run the 143–144 GiB DeepSeek-V4-Flash-0731 UD-Q4_K_XL GGUF on four RTX 3060 12GB cards while keeping a 360k–376k context window. Hardware…
submitted by /u/johnnyApplePRNG [link] [comments]
submitted by /u/anderspitman [link] [comments]
submitted by /u/johnnyApplePRNG [link] [comments]
Apparently a second version of DFlash from the original authors of DFlash GGUF quants are already made available with an accompanying llama.cpp PR: ht…
Who needs GPUs? submitted by /u/DeltaSqueezer [link] [comments]