arXiv cs.CLSeptember 24, 2026
Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding
Excerpt
arXiv:2609.27678v1 Announce Type: new Abstract: Contract inference requires multiple judgments about a shared document, but aggregate accuracy can conceal changes in the individual decisions. Repeated agreement is also insufficient: a model may consistently return the wrong answer. In this paper, we compare Jev with nine language models on ContractNLI, evaluating inference cost, response time, average correctness, and correctness across repeated request conditions. Controlled comparisons vary hy