Research questionHow can LLM prompts be automatically refined from recurring reasoning errors without laborious manual engineering?Prompt performance can depend heavily on wording and instruction order, making manual refinement costly. Methods that inspect only individual examples or small batches may fail to identify recurring errors and make targeted corrections.