Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
We just released the first version of our Creative Writing benchmark, comparing 24 LLMs against human writers across 475 writing prompts. Creative wri…
submitted by /u/Thatisverytrue54321 [link] [comments]
According to people familiar with the matter, Naive AI, an AI startup founded in February this year by Tsinghua University professor Dai Jifeng, has c…
Hey r/LocalLLaMA , Prism-LM recently released its Bonsai 2 QAT models based on Qwen3.8, and they quickly gained traction. In our evaluation, the model…
MiniMax has open-sourced the terminal version of MiniMax Code: https://github.com/MiniMax-AI/minimax-code How can developers verify the content that e…
do you want some omni? here is omni for you 1. 🧭 Overview This repository hosts two checkpoints of the Realtime-Venus system: Realtime-Venus-Omni ( R…
Yesterday I saw that prismml just dropped the new bonsai which is a heavily-quantized version of qwen 3.8 27b claiming over 98% top-1% comparing to fp…
I have always posted about budget builds on here, and often asked how we are going to run the next big models. Often Plenty of downvotes too or folks…
I’m thinking about Bonsai 2… submitted by /u/JLeonsarmiento [link] [comments]
https://preview.redd.it/bqjrvgknt3qh1.png?width=2810&format=png&auto=webp&s=ec9f25709d6f2951a530228d430e0b8ecdbd8738 source super interesting read fro…
submitted by /u/Thrumpwart [link] [comments]
Everyone now talks about the architecture that's not auto regressive and does lightning fast probability prediction with a json schema. I worked on th…
Note not my work, but something i found and wanted to share so hopefully more people can push this along even further. https://github.com/nasone32/lla…
I had Astra run a multi-hour investigation into speeding up local MoE inference while keeping the model weights and quantization unchanged. The experi…
It's not the brightest bulb but it's my new drudgework model for read+find or code tasks I'm willing to let it brute force. Even if it takes 20x more…
Hey all, Henry from Cactus Compute here, I kinda wanted to share our latest model and get feedback from the family :) Needle 3 is a small foundation m…
Been building this for a few months, mostly for myself, and it just got a proper release so figured I'd post it. It's a native GGUF inference runtime…
Great news! AMD is also considering accepting payment in organs! Slightly less sarcastically, grab what you can, while you can. Waiting is becoming ve…
CPU : epyc 7262 MB : ROMED8-2T GPU : V100 16GB PCIE x6 I have built a server with six V100 GPUs. I am now testing it and plan to eventually run Qwen3.…
My understanding so far: You take an LLM and use it without thinking (That's what openjev does?) You leave out the text generation in the end and take…