Research questionHow can language models reliably follow instructions containing many simultaneous constraints?As prompts accumulate explicit requirements, language models often satisfy some while overlooking others. Existing datasets and benchmarks generally stop at ten constraints, making performance on denser real-world instructions difficult to train for and measure.