arXiv cs.AIOctober 7, 2026
Evolving in Thought Space: Training a Small Model at Test Time Unlocks Better Discoveries
Excerpt
arXiv:2610.06269v1 Announce Type: new Abstract: Open-ended scientific discovery often requires repeatedly proposing and evaluating candidate solutions. LLM-based systems can support this process by generating and refining executable solutions from verifier feedback. Methods such as TTT-Discover use test-time training (TTT) to update the solution-generating LLM from verifier feedback, adapting its generation policy to improve subsequent proposals on the target problem. However, this becomes expen