Research questionHow can we tell whether LLMs follow coherent, human-like prerequisite relationships in mathematical reasoning?A model can answer individual mathematics questions correctly while violating expected prerequisite dependencies or failing to use related knowledge in context. These structural inconsistencies may remain hidden by accuracy-based and LLM-as-judge evaluations.