Research questionCan prompt phrasing reliably improve LLM-derived chemical features for drug-toxicity prediction?Minor changes in prompt phrasing can alter LLM outputs, making it unclear whether prompt optimization produces stable chemical features for toxicity models. This variability complicates the use of LLM-generated features in a costly drug-development process.