arXiv cs.AIOctober 7, 2026
$\pi^2$: Structure-Originated Reasoning Data Improves Long-Context Reasoning Ability of Large Language Models
Excerpt
arXiv:2604.05114v2 Announce Type: replace-cross Abstract: We study a QA curation pipeline for improving long-context complex reasoning in large language models (LLMs). Our approach, $\pi^2$, constructs high-quality reasoning data through rigorous QA curation: 1) extracting and expanding tables from Wikipedia, 2) from the collected tables together with relevant metadata, generating complex reasoning questions whose answers are automatically determined and validated through dual-path code executio