Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
OpenAI has decided to fully shut down a protein design project I'm working on for a client. Needless to say, open weight models are the only way forwa…
People keep on getting confused about this, so I looked at the safetensors on hf. The title should have been "Deepseek V4.1 Flash is 748B total/552B b…
Hoping to see smartest medium size models soon & later with all available optimizations/architectures/etc.,. Thanks Deepseek! Ex 1: 30-50B MOE + 10-15…
submitted by /u/t4a8945 [link] [comments]
Here we go again, DeepSeek is back again with a new model V4-1 Flash A multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and sup…
In our chat app "harness" we recursively generate summaries, L1 → L2 → L3. L1 summaries are more factual extraction than coherence, then get rolled in…
I’ve been using Qwen a lot recently and I’m honestly very impressed with it. For some of my tasks, Qwen seems to understand what I’m actually trying t…
Hi, I'm wondering what settings you are using in order to run Qwen3.8-Flash-Next on your devices? I'm especially interested in setups with 96GB VRAM.…
(hopefully this is okay to here - it seems like audio models and image / video modeals is allowed but yeah this is a bit different) So I've been doing…
I was inspired by Bijan Bowen video - Subway FPS https://youtu.be/6kjXzTVmT58?t=1035 Wondered how far I can push Qwen 3.8 27b so I used a plan made by…
I don't think anyone posted about this here, but Qwen released a finetuned version of 3.5 4 for driving. The full Bf16 checkpoint is 9B. This is a ver…
https://www.reuters.com/technology/us-accuses-chinese-ai-firms-industrial-scale-theft-ai-technology-2026-09-08/ https://www.nbcnews.com/tech/tech-news…
The fix: run the server detached/headless instead of keeping it attached to a console window. RTX 5090, ~27B NVFP4 model via ninfer: Terminal focused:…
new model of nex benchmark seems like good https://huggingface.co/nex-agi/Nex-N2.5-Max submitted by /u/Lordaizen639 [link] [comments]
Running the full deepseek-ai/DeepSeek-V4-Flash-Vision-Exp on consumer Ampere — 10-12x RTX 3090, SM86-compatible vLLM build. 285B MoE, FP4 experts + FP…
Server rebuild to custom loop. Temps on the gpus went from upper 80s to mid 40s under load, 30c idle. 5950X 64gb ddr4 2x rtx titans (24gb vram ea) 1x…
Privacy have been concern of many of us to have their own hardware to run llms, and here's another reason why: two mathematicians spent a year crackin…
submitted by /u/Arcuru [link] [comments]
Hi, I'm the co-creator of this Qwen3.8-Flash-Next engine support in MLX-serve. I've been tuning this one to run both fast, efficient and correct up 1m…