arXiv cs.AIOctober 7, 2026
Knowing the Rules, Applying the Rules: Evaluating Language Models on Traditional Chinese Bazi
Excerpt
arXiv:2610.05682v1 Announce Type: cross Abstract: Knowing domain rules does not guarantee applying them to a case. We study this distinction in traditional Chinese Bazi through 3,000 Chinese multiple-choice questions spanning 14 Theory and 11 Case categories. Six endpoint systems are evaluated, with primary results reported on a 2,492-item model-informed refinement. Theory accuracy exceeds Case accuracy for every system, and gaps of 16.60-29.56 percentage points remain when invalid responses are