arXiv cs.CLAugust 17, 2026
Lost in Historical Time? A Polish History Matura Benchmark for Large Language Models
Excerpt
arXiv:2608.12343v2 Announce Type: replace Abstract: Language models are widely used by students as knowledge sources, yet benchmarks rarely assess their interpretative historical reasoning. We evaluate eight leading LLMs on the Polish high school exit exam (Matura) in history - three official papers from 2023-2025, comprising short-answer questions and extended essays - and compare model performance against the human examinee population. Although models score near the ceiling, aggregate scores m