
An evaluation of whether stating requirements and output constraints in the initial prompt reduces correction exchanges and tokens per solved task. The study runs 14 machine-checkable tasks across lazy and front-loaded prompt arms, thinking off and default adaptive thinking, with five repetitions per cell and one verifier-guided correction turn after failures. Across 280 first turns and 120 corrections, front-loaded prompts used 45,463 recorded tokens versus 160,799 for lazy prompts, while solving 134 of 140 repetitions versus 96 of 140. The paper reports the method, verifier design, task-category results, counterexamples, limitations, and reproducibility artifacts.
