← Back to all articles
Reddit r/MachineLearningSeptember 21, 2026

Jev's calibration was measured. The LLMs won [D]

Excerpt

Source: Jev Benchmarks Its training method is literally called "Reinforcement Learning for Calibrated Decisions." Calibration gap vs human labels (lower = better): Yes/no: Jev 5.0, Gemini 3.8 Flash 2.0 Pick-one: Jev 9.8, DeepSeek V4.1 Flash 2.8 Rubric: Jev 19.7, GLM-5.3 12.9 It held to 95% accuracy, Jev still handles more decisions alone than any of them (86% of yes/no). Worse calibrated, better at knowing when it's right. submitted by /u/frappuccinoCoin [link] [comments]