← Back to all articles
arXiv cs.AIAugust 18, 2026

ALPS: Measuring Valid Creativity in Large Language Models with Mathematical Construction

Excerpt

arXiv:2608.15979v1 Announce Type: new Abstract: Large language models produce outputs presented as discoveries - new proofs, conjectures, or molecules. Whether such an output that appears creative is truly original and effective is hard to establish: open-ended outputs require subjective judgment, the output may replicate something seen in training, or the task may be too simple to need creativity. We present ALPS (Austin-Law Proof-Synthesis), a benchmark that designs a task to measure valid cre