Research questionHow can we quantify and reduce divergent, nonsensical reasoning in large language models?LLMs can generate multiple incompatible reasoning chains from the same prompt, including branches that become implausible or nonsensical. This makes it difficult to distinguish productive exploration from reasoning instability and to reduce the latter.