Latest AI/ML News

770 articles · Reddit r/LocalLLaMA

Reddit r/LocalLLaMASep 3, 2026

ik_llama.cpp merged qwen4exp MTP support yesterday (PR #2369, mine, reviewed and tested by four other people on their own hardware). It's on main now,…

Reddit r/LocalLLaMASep 3, 2026

124B total parameters, 5.1B activated parameters, and a 256K context window submitted by /u/Bestlife73 [link] [comments]

Reddit r/LocalLLaMASep 3, 2026

I liked the Nvidia that focused on just GPUs for gaming, not on the Nvidia of today which seem want power consolidation. Modelscope is another platfor…

Reddit r/LocalLLaMASep 3, 2026

How was your experience with it and is it just benchmaxed or is it really that good? submitted by /u/Personal-Try2776 [link] [comments]

Reddit r/LocalLLaMASep 3, 2026

Looking into the new Qwen architecture, I was curious if you could modify the Ngram PLE Table to make it work like a long-term knowledge database. It…

Reddit r/LocalLLaMASep 3, 2026

more sizes (probably still uploading): https://huggingface.co/IFM/K2-Horizon-32B-GGUF https://huggingface.co/IFM/K2-Horizon-7B-GGUF https://huggingfac…

Reddit r/LocalLLaMASep 3, 2026

ChatGPT is down r/ChatGPT Claude is down r/ClaudeCode Grok is down r/grok my local llama.cpp works as always submitted by /u/jacek2023 [link] [comment…

Reddit r/LocalLLaMASep 3, 2026

submitted by /u/Few_Painter_5588 [link] [comments]

Reddit r/LocalLLaMASep 3, 2026

submitted by /u/mailto_devnull [link] [comments]

Reddit r/LocalLLaMASep 3, 2026

submitted by /u/SarcasticBaka [link] [comments]

Reddit r/LocalLLaMASep 3, 2026

One of the biggest unlocks was getting my speech to text functioning reliably, especialyl for coding. Right now I'm at the point where when I have to…

Reddit r/LocalLLaMASep 3, 2026

Obviously more memory is good, more context, bigger models, but some jumps don't actually unlock a meaningful difference in ability to run different o…

Reddit r/LocalLLaMASep 3, 2026

Disclaimer: I'm the builder. --- Over the past 2 yrs, I was working on a SaaS ML project & got very interested in ML/DL/AI. As everybody else, there w…

Reddit r/LocalLLaMASep 3, 2026

I've been building Arcon around Qwen3-4B + LoRA. Instead of just making it a chatbot, I'm experimenting with persistent memory, personality/mood, inte…

Reddit r/LocalLLaMASep 3, 2026

75B MoE is an interesting size to check, you can run it today (no MTP support yet) The model employs a hybrid MoE architecture with interleaved Mamba,…

Reddit r/LocalLLaMASep 3, 2026

Sharing some numbers because most posts on this model are either using a single 3090 or unified systems from what I've seen. My current setup: 2x RTX…

Reddit r/LocalLLaMASep 3, 2026

This is mostly for setting up for expectation, since personally without LLM i could take 3 days (15 hours of active programming) to debug or implement…

Reddit r/LocalLLaMASep 3, 2026

***This screenshot is an app i made for creating a dataset from scratch, this is not a real chat with the model*** First of all i want to shout out ev…

Reddit r/LocalLLaMASep 3, 2026

Is anyone interested in my Chrome browser add-on that uses local LLMs to move thousands of unsorted bookmarks into a smart list of automatically calcu…