Research questionHow can black-box LLMs resist jailbreaks without weight access or retraining while preserving benign-query utility?Jailbreak prompts can bypass safety alignment, while defenses that depend on model weights or internal signals cannot be added to many deployed systems. The practical difficulty is reducing malicious compliance without disrupting ordinary user requests.