Research questionHow can audio-captioning datasets represent fine-grained acoustic detail and perceptual ambiguity for better audio retrieval?Many audio-captioning datasets provide generic descriptions and only one caption per clip, even though listeners may describe the same sounds in different valid ways. Missing acoustic detail and semantic variation can limit models trained for audio retrieval and related audio-language tasks.