Latest AI/ML News
2649 articles · 👥 Community Buzz
Source: Jev Benchmarks Its training method is literally called "Reinforcement Learning for Calibrated Decisions." Calibration gap vs human labels (low…
Small, free finding. I use local Qwen models for typed decisions: a state plus a question with fixed allowed answers, and I read the probability of ea…
# the What An engine to run Gemma 4 31B on blackwell under massive concurrency and rather specific workload patterns. I've been waiting for someone to…
https://preview.redd.it/nulsv53o8vqh1.png?width=4500&format=png&auto=webp&s=74765dbd409f4c221640f9f6000a685f6fdbb242 Spent weekend benchmarking the Sp…
I decided to do this experiment to see what would be the ability of an LLM model to, in a single shot, transfer the knowledge it may or may not have a…
Since there’s no comparison chart on the model page, I asked Perplexity to compare it against some relatively small open-weight models in a similar si…
Hey all! I am Aritra from Hugging Face. I wanted to share an update on the `tokenizers` library that we have at Hugging Face. It has gone under major…
submitted by /u/Bitter-College8786 [link] [comments]
we're so back?!? submitted by /u/VoiceApprehensive893 [link] [comments]
Isn't this what simple neural networks have been able to do for years? Doesn't seem anything special to me. submitted by /u/Manerfish [link] [comments…
Quote: DeepSeek is training a 2T-parameter model and plans to eventually build an 8T-parameter model. https://x.com/wallstengine/status/21019828436563…
submitted by /u/Bestlife73 [link] [comments]
I’ve been working on making small models more capable at agentic coding and work, because most people in the world don’t have the sort of hardware nee…
submitted by /u/fallingdowndizzyvr [link] [comments]
submitted by /u/themixtergames [link] [comments]
https://huggingface.co/yandex/AliceAI-Foundation-80B-A3B-Base It's not a Qwen3 finetune, it's actually its own fully custom architecture. No Llama.cpp…