Get Started
Home
Topics
Search
Library
Research questionHow can audio-captioning datasets represent fine-grained acoustic detail and perceptual ambiguity for better audio retrieval?Many audio-captioning datasets provide generic descriptions and only one caption per clip, even though listeners may describe the same sounds in different valid ways. Missing acoustic detail and semantic variation can limit models trained for audio retrieval and related audio-language tasks.
Audio & Speech
Audio & Speech Processing
Evaluation & Benchmarks
Information Retrieval
Machine Learning
Multimodal Models
Sound
Latest papersRecent research connected to this question, newest first.SonicCaps: Large-Scale Diverse and Fine-Grained Captioning for Improved Audio-RetrievalSonicCaps contains about 15 million captions paired with about 700,000 audio clips, with roughly 24 captions per clip covering descriptions, rephrasings, stylistic or verbosity variants, and semantic tags. The captions were generated by Qwen3-Omni using audio and text, then assessed by humans; training CLAP models with the dataset was reported to improve audio retrieval and zero-shot classification across public and commercial benchmarks.research paper · Sep 2, 2026
Related questions
How can image-text retrieval focus on caption-described attributes while ignoring unmentioned visual information?How can fine-grained audio-visual segmentation learn new classes continually without semantic drift or co-occurrence confusion?How can single-pass image captioning capture fine-grained visual details without multi-stage latency?How can discrete audio tokenizers preserve semantics and acoustic fidelity for both understanding and generation?