Research questionHow can safety-tuned language models distinguish harmful requests from benign ones with risky wording?Safety tuning can cause models to rely on risky-looking words rather than the request’s actual intent. As a result, harmless users may be refused even when the model should respond helpfully.