Research questionHow can black-box language models reliably follow procedural instructions at inference time for downstream trajectory repair?Instruction-following failures can leave downstream components without the procedural steps needed to inspect or repair a generated trajectory. Better procedural compliance may not improve final-answer accuracy and can change how early the model commits to an answer.