Research questionHow should open reasoning language models be selected under prompt and deployment-resource trade-offs?Prompting can change the relative performance of reasoning models, making isolated accuracy comparisons hard to interpret. Latency and memory constraints further affect which model is suitable for deployment.