Latest AI/ML News
2649 articles · 👥 Community Buzz
Retrieval benchmarks sometimes feel benchmaxxed by models, so we wanted to find a way to tie it as close as possible to my objective: finding the arti…
That's the whole question really. Which uncensored / abliterated versions of Qwen Next Flash have you used, and what's your experience? Pros / cons? s…
Github: https://github.com/OpenEuroLLM/ComplexKDA HuggingFace: https://huggingface.co/collections/openeurollm/complexkda Arxiv: https://arxiv.org/abs/…
A note: this entire thing is 100% human written, not even AI drafted or edited. So, enjoy. Or not. edit for a tldr that completed just after posting:…
I didn't know my set up was outperforming nearly everyone until reading another discussion where people were struggling getting half of that speed wit…
https://preview.redd.it/ybwqqbrst0rh1.png?width=811&format=png&auto=webp&s=1e40c5b53e304caf2a10efa2baa6f998b9a9a0bb MiMo-V2.6-Distill-Qwen-9B is a 9B…
submitted by /u/Aggravating-Push-207 [link] [comments]
Basically the title.We did not get a new moe model with qwen 3.8 and Alibaba did not announce any small moe models on apsara.I know we might get an an…
submitted by /u/johnnyApplePRNG [link] [comments]
This post is written by a human and I'd appreciate it if you treated it as such. Thanks. So, I've been noticing a pretty clear interest in developing…
TLDR: We demonstrate and explain the difference in expressivity of Gated Deltanet (GDN) and Kimi Delta Attention (KDA). We show how the full diagonal…