Latest AI/ML News
2649 articles · 👥 Community Buzz
Hi, I have a solo paper on arxiv from my masters studies, which was my part-time work and a few months back I got back to it and tried to finally make…
The total training cost was just $3.5M. The model comes with a live benchmaxxing dashboard. https://preview.redd.it/89uurv5r21rh1.png?width=1518&forma…
Although it makes mistakes and use more tokens, but after some corrections and steering , It(max) gives really good outputs like on par with 5.6 sol a…
Recently, the new DeepSeek-V4.1-Flash architecture showed how a causal encoder-decoder can work, but it was trained from scratch. Model Grafting does…
Imagine Git for model fine-tunes that also saves you storage. DeltaTensors compresses fine-tuned model checkpoints by storing the weight difference fr…
I decided to see what effects recent PRs have had on the performance of the two models I care about, DSv4 Flash 0731 and Qwen 3.8 Flash Next, on my ha…
This experiment is live, you can inspect all the internal reasoning, memories, attempts here: https://artificium-covering-experiment.gr.bio/ The probl…
Been playing with judge models in my eval pipeline, so this landed at the right time. Kev is a small family of Jev-architecture decision models (0.8B,…
https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B That's serious. submitted by /u/Beamsters [link] [comments]
I shipped something I've been building for the last few weeks : phantom-kv , a refusal-removal system for large language models that doesn't touch a s…
submitted by /u/tengo_harambe [link] [comments]
https://preview.redd.it/bpbc9i6hizqh1.png?width=1270&format=png&auto=webp&s=e8aa8301895735a05c3c61a5e793018a23d1cac5 I wanted to share a quick update:…