← Back to all articles
Reddit r/MachineLearningSeptember 28, 2026

Reduced my Jev judge’s calibration error [D]

Excerpt

I benchmarked Jev on TRIVIA+ dataset using an untouched 645-example test set. Before calibration: ECE: 0.0982 After learning from human-labelled examples: ECE: 0.0313 That’s a 68.1% reduction in calibration error . But hallucination-detection F1 only moved: 0.5833 → 0.5877 So what improved? Not the judge’s ability to classify examples. Its confidence became much more aligned with reality . And that distinction matters. If a judge score is only being displayed on a dashboard, maybe not much. But