Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
ik_llama.cpp merged qwen4exp MTP support yesterday (PR #2369, mine, reviewed and tested by four other people on their own hardware). It's on main now,…
124B total parameters, 5.1B activated parameters, and a 256K context window submitted by /u/Bestlife73 [link] [comments]
I liked the Nvidia that focused on just GPUs for gaming, not on the Nvidia of today which seem want power consolidation. Modelscope is another platfor…
How was your experience with it and is it just benchmaxed or is it really that good? submitted by /u/Personal-Try2776 [link] [comments]
Looking into the new Qwen architecture, I was curious if you could modify the Ngram PLE Table to make it work like a long-term knowledge database. It…
more sizes (probably still uploading): https://huggingface.co/IFM/K2-Horizon-32B-GGUF https://huggingface.co/IFM/K2-Horizon-7B-GGUF https://huggingfac…
ChatGPT is down r/ChatGPT Claude is down r/ClaudeCode Grok is down r/grok my local llama.cpp works as always submitted by /u/jacek2023 [link] [comment…
submitted by /u/Few_Painter_5588 [link] [comments]
submitted by /u/mailto_devnull [link] [comments]
submitted by /u/SarcasticBaka [link] [comments]
One of the biggest unlocks was getting my speech to text functioning reliably, especialyl for coding. Right now I'm at the point where when I have to…
submitted by /u/comperr [link] [comments]
Obviously more memory is good, more context, bigger models, but some jumps don't actually unlock a meaningful difference in ability to run different o…
Disclaimer: I'm the builder. --- Over the past 2 yrs, I was working on a SaaS ML project & got very interested in ML/DL/AI. As everybody else, there w…
I've been building Arcon around Qwen3-4B + LoRA. Instead of just making it a chatbot, I'm experimenting with persistent memory, personality/mood, inte…
75B MoE is an interesting size to check, you can run it today (no MTP support yet) The model employs a hybrid MoE architecture with interleaved Mamba,…
Sharing some numbers because most posts on this model are either using a single 3090 or unified systems from what I've seen. My current setup: 2x RTX…
This is mostly for setting up for expectation, since personally without LLM i could take 3 days (15 hours of active programming) to debug or implement…
***This screenshot is an app i made for creating a dataset from scratch, this is not a real chat with the model*** First of all i want to shout out ev…
Is anyone interested in my Chrome browser add-on that uses local LLMs to move thousands of unsorted bookmarks into a smart list of automatically calcu…