Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
This post is written by a human and I'd appreciate it if you treated it as such. Thanks. So, I've been noticing a pretty clear interest in developing…
Although it makes mistakes and use more tokens, but after some corrections and steering , It(max) gives really good outputs like on par with 5.6 sol a…
Recently, the new DeepSeek-V4.1-Flash architecture showed how a causal encoder-decoder can work, but it was trained from scratch. Model Grafting does…
Imagine Git for model fine-tunes that also saves you storage. DeltaTensors compresses fine-tuned model checkpoints by storing the weight difference fr…
I decided to see what effects recent PRs have had on the performance of the two models I care about, DSv4 Flash 0731 and Qwen 3.8 Flash Next, on my ha…
This experiment is live, you can inspect all the internal reasoning, memories, attempts here: https://artificium-covering-experiment.gr.bio/ The probl…
Been playing with judge models in my eval pipeline, so this landed at the right time. Kev is a small family of Jev-architecture decision models (0.8B,…
https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B That's serious. submitted by /u/Beamsters [link] [comments]
I shipped something I've been building for the last few weeks : phantom-kv , a refusal-removal system for large language models that doesn't touch a s…
submitted by /u/tengo_harambe [link] [comments]
https://preview.redd.it/bpbc9i6hizqh1.png?width=1270&format=png&auto=webp&s=e8aa8301895735a05c3c61a5e793018a23d1cac5 I wanted to share a quick update:…
Small, free finding. I use local Qwen models for typed decisions: a state plus a question with fixed allowed answers, and I read the probability of ea…
# the What An engine to run Gemma 4 31B on blackwell under massive concurrency and rather specific workload patterns. I've been waiting for someone to…
https://preview.redd.it/nulsv53o8vqh1.png?width=4500&format=png&auto=webp&s=74765dbd409f4c221640f9f6000a685f6fdbb242 Spent weekend benchmarking the Sp…
I decided to do this experiment to see what would be the ability of an LLM model to, in a single shot, transfer the knowledge it may or may not have a…
Since there’s no comparison chart on the model page, I asked Perplexity to compare it against some relatively small open-weight models in a similar si…
Hey all! I am Aritra from Hugging Face. I wanted to share an update on the `tokenizers` library that we have at Hugging Face. It has gone under major…
submitted by /u/Bitter-College8786 [link] [comments]
we're so back?!? submitted by /u/VoiceApprehensive893 [link] [comments]
Isn't this what simple neural networks have been able to do for years? Doesn't seem anything special to me. submitted by /u/Manerfish [link] [comments…