Research questionHow can prompt representations distinguish specificity to expose fine-grained LLM weaknesses?Prompts about the same topic can differ substantially in how specific or difficult they are. Standard vector representations may place them too close together, obscuring which kinds of instructions reveal model weaknesses.