Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
submitted by /u/infieldmitt [link] [comments]
I finished my home inference server. First I tried Lenovo p620 workstation and while it’s a good value overall it pissed me off with a ton of propriet…
I have a full M5 Pro Macbook Pro with 48GB of RAM, I'm just getting into this local space. Like many of you, the costs of using frontier/cloud models…
A few people here mentioned interest in a way to test their custom Pi setups, so I figured I’d drop this here: RoastMyHarness The basic idea is a smal…
I have been battling with my Qwen3.8:27b setup on my rtx 5080 16gb. I am using llama.cpp to run a nvfp4 version of qwen3.8:27b llama-b10699-bin-win-cu…
It’s always bothered me that after fine-tuning a model for a project, there isn’t a particularly easy way to host it without either running it locally…
While browsing a benchmark list site, I spotted a recently published 33B parameter model which claimed to beat Qwen3.8 27b on the ArtificialAnalysis (…
As you all know, Qwen3.8 Flash Next on mainline llama.cpp is still in a pretty experimental stage, but a lot of community forks are trying to get it t…
New website to download models in case HF starts censoring or limiting access. submitted by /u/Thrumpwart [link] [comments]
submitted by /u/One-Replacement-37 [link] [comments]
For those using these models for coding in larger projects where things can get complex, do you find yourself using the 8-bit quants if you have enoug…
PP improvements for RDNA2(MI50, MI60 are included in benchmarks). Check bottom comments of PR to see updated pp t/s stats. submitted by /u/pmttyji [li…
Kimi routed some PLA requests to Claude for distillation purposes without warning the PLA users. There is rumor that 16 Moonshot employees were arrest…
Hey everyone — I’m building TensorSharp , an open-source LLM inference engine. Here are the latest DeepSeek V4.1 Flash GGUF results using its native g…
Q2 is there and Q4 is uploading as I type. Has his github been updated yet? How do you run this? https://huggingface.co/antirez/deepseek-v4.1-flash-gg…
What I have: - CPU: EPYC 7551 (32c/64T, Zen 1) - Board: Supermicro H11SSL-i (SP3), Rev 2.0 - RAM: 128 GB DDR4-2133 (all 8 channels full) - GPU: 2x RTX…
Another new model dropped in the course of this week that is well deployable on consumer hardware: Nex N2.5 Mini I went with the recommended settings…
Reports of the demise of coders may have been exaggerated. submitted by /u/SteppenAxolotl [link] [comments]
Hola all. Do you guys mind sharing your LLama.cpp config and system setup details for Qwen3.8 Flash Next? Model's quite big and tryining many combinat…
I have enabled the vision for the CIRU Strix UL4 quant of Qwen 3.8 flash next (others quants likely perform very similar) and tried it on a few things…