Latest AI/ML News
770 articles · Reddit r/LocalLLaMA
Hi All, I am currently contemplating an upgrade from my 5060ti 16gb. I am getting ~40t/s with 130k context on Qwen 3.8 IQ3_S HF quant. I am running ll…
Prompt Flow-> Test to Image -> Image to Video - Stitch" src="https://external-preview.redd.it/NDRseWJsd2o5a21oMb0OE-fBS75cTqn-cbhZROlP23RQAQcyt0HE8j_n…
submitted by /u/Fcking_Chuck [link] [comments]
In 2025, they were planning to release an iphone with mobile hbm in 2027; perhaps they have scraped this idea due to higher memory prices. It will be…
We used HFlow to evaluate the latest open weights VLMs for processing egocentric data. This was based on Build AI's Egocentric-10k evaluation , which…
Qwen 3.8 Flash Next (80gb) now at 3.5 tok/s on 12gb mid range android phone thanks to some optimizations and with a low quantization on dense part. I…
Been working really hard for the past month to bring to the community all these models, the hardest was for sure LongCat-Flash-Lite-Sparse who require…
Hello everyone, Just sharing this cool and useful feature in the KernelAI mobille app. It works on pretty much any model, and the ability to run both…
Since I spent the time to figure it out and it is not like it will make me any money ever. I think I'd share with you all what I managed to cobble tog…
This might be worth it for some small business. 7.1tb vram bandwidth submitted by /u/SpendLucky1273 [link] [comments]
So, ideally for this thread we exclude the ones that everyone on here is already well aware of and discussing on here a lot, like N-gram, quantization…
I find both Qwen 3.8 27b and Qwen 3.8 Flash Next difficult to read. Here's some examples of what I mean: **Model-visible tool set per turn** (assemble…
Just noticed this on the website. At their current price tiers for the memory SKUs (32, 64, 128) I'd expect this to be ~ 4.5k for the motherboard. The…
OpenClaw was all the rage a few months ago but the hype seems to have died down. Are you guys using it for any of your needs? submitted by /u/cdrfrk […
submitted by /u/liright [link] [comments]
audio.cpp 0.7 is out :) This release adds a lot of new audio models and a new way to compare them locally. Audio.cpp is now at 62 model families and 8…
ok bit more context: it's actually a QAT Q2 for Qwen 3.8 27 B: https://huggingface.co/sdkyuan/qwen3.8-27B-qat-q2_0-gguf QAT Q2 for DFlash model: https…
Everything has changed in two months: DS4 0731 flash was the start of a wave that is taking open weights to paradise. It is easy to think that Qwen 4…
submitted by /u/DjCanalex [link] [comments]