arXiv cs.LGOctober 7, 2026
A Systematic Study of Small Language Models on Abstract Reasoning Tasks
Excerpt
arXiv:2610.08680v1 Announce Type: new Abstract: Endpoint accuracy on abstract-reasoning benchmarks does not reveal whether a language model has acquired a transferable rule or fit distribution-specific regularities. We study this distinction in small language models on the ARC-TGI benchmark, which organizes abstract grid transformations into controllable task families and supports resampling, spatial shifts, and cross-benchmark transfer. Across more than 1,000 runs, we profile decoder-only, enco