Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
Im looking to set up a new local llm (probably on unsloth studio as that seemed to be doing pretty well last time I tested it). This one won't need to…
Look at me: I am the frontier Lab now Huihui-Qwen3.6-35B prompt (on Pi): "In the folder u/source/ you will find 6 files, text files, that are commonly…
I get why they exist and in almost mostly any other hardware field we can see clearly the difference and what it respects throughout, but with ai, its…
I have a gigabyte ds3h v2 b450 motherboard. Currently hosting an RTX 3090 I have a spare 3070 and I wondered, can I run both? Got myself a riser cable…
I love the Ornith 35B local models, 1.0 has been running my HAM radio rig for me. I have a hackRF receiver and a 5 watt quansheng portable the both ru…
Hello, I have a few questions that I can't seem to find a clear answer to. Does it make sense to make your own GGUF? I noticed that when I compile lla…
For example: Instead of saying: "I created this new ID" It says: "I minted this new ID" Instead of: "This alternative path is available" It says: "thi…
This is a genuine community of real generally respectful adult human beings. Despite the enthusiasm all of you have for local AI, you can recognize th…
I was experimenting with deepseek harness when found that even if you don't use deepseek models, you can configure the web_search tool with their api…
Using llama.cpp I seem to be unable to get my to GPUs working tougether correclty, so I need help somehow. Setup: 96GB RAM, one Blackwell 5000 (48GB)…
submitted by /u/johnnyApplePRNG [link] [comments]
New/Old benchmark that provides a lot of answers for local LLM. I present to you a new test that I developed somewhat by accident: https://huggingface…
Idea is simple-ish in abstract: instead of running normal prefill over all of a given prompt, split it into parts - generates caches for part A and pa…
I was finally able to replicate tensor level allocation outside the Gemma family. https://huggingface.co/ByteOtter/Qwen3.5-4B-CADA-IQ2_XS After the Ge…
No GGUFs yet on Huggingface though. submitted by /u/Ihtien [link] [comments]
A new project was released yesterday and I have the opportunity to test it today. Papper: https://arxiv.org/abs/2608.16157 Github: https://github.com/…
https://preview.redd.it/84zi5nsdawkh1.png?width=2368&format=png&auto=webp&s=1109e69db807b153064b1f5b61d22cf1e9fbca05 Another user posted the benchmark…
submitted by /u/DeltaSqueezer [link] [comments]
Has anyone else experimented with ubatch size when running GLM-5.2 locally? I was testing the 226 GiB GLM-5.2-UD-IQ2\_XXS GGUF on 3x RTX PRO 6000 Blac…
You can find the changelog and source code here: https://github.com/ggml-org/llama.cpp/releases/tag/v0.2.0 Associated pre-build is here: https://githu…