Latest AI/ML News
2649 articles · 👥 Community Buzz
edit: confound Hello all, I am working on a object classification with a automotive radar point clouds. I compared many models and feature vectors. On…
I've been out of the loop for some time. Is there already an official way to offload ngram to RAM or SSD in something like Unsloth Studio? Interested…
I’ve been building "Spomin", a router that replaces context with summaries directly in the KV cache. The goal is to keep long running sessions going w…
Developer's own thread: https://www.reddit.com/r/LocalLLaMA/s/adp1cGZZe9 Code: https://github.com/Inovello/llama.cpp/tree/flashnext-e06 My hardware: 2…
I wanted to see how my fully local home voice assistant compared to the latest GPT Live, so I tested it using the same conversation used in their "Imp…
A huge favor to ask llama maintainers - please implement this feature. Even with one 3090 card there will be tangible improvements in decode speed on…
Can you still do something with a 2050 or something like it? I mean for office work, loading embedding, reranking and chat models not at the same time…
Github Repo. Blog post. 💡 TL;DR (from the Github Readme) Spend less without making the agent do less useful work. SoL-Pi is a standalone extension fo…
Like many of you, I have seen many posts and tweets in the last weeks complaining about Artificial Analysis being "broken", "meaningless", and "bought…
Astra's launch has produced a strange discourse. The reporting that broke the story framed the model's use of recurrent depth primarily as a safety re…
Hello Guys, Just sharing this harness I've created, this runs entirely on chromium based browser. No installation is required, it has its limitations…
Planning to build a PC mainly for local LLMs/coding agents. I keep seeing 3090 + Qwen 3.8 Flash Next benchmarks, but could not find enough info for th…
submitted by /u/TGSCrust [link] [comments]
Hi! I've been following this community for quite a while and have difficulty figuring out what to put on my 3080 12gb - I know Qwen 3.6 35B 3A was the…
Nice pp improvements for RDNA4(R9700) & 3.5(RX 9060 XT, 8060S). More good numbers on large context. PR has detailed benchmarks. u/ilintar 👍 submitted…