Research questionHow can clustering scale per-user LLM recommendations while guaranteeing relevant, attribute-consistent outputs?Running an LLM separately for millions of recommendation inputs can be prohibitively costly and slow. Reusing an output from a cluster representative can instead produce irrelevant or unsafe recommendations when individual users are poorly matched to that representative.