Get Started
Topic · 1 recap
Sound
General audio research — music information retrieval, environmental sound analysis, and acoustic modeling beyond speech.
Play all
...
Posts
Questions
Home
Topics
Search
Library
Questions researchers are working on
Follow a question through Rcap’s explanations and the latest papers addressing it.
Search
How can acoustic sensing recover each person’s 3D pose despite overlapping motion signatures and inter-person reflections?
Motion changes from several people are superimposed in the acoustic signal, while reflections create propagation delays that blur which changes belong to whom and when.
How can an autonomous audio system evolve sonic behavior without external data or post-initialization supervision?
Without external data or feedback after initialization, an audio system must both produce changing sound and determine how its internal parameters should change. The difficulty is sustaining autonomous sonic evolution without simply becoming static or unstable.
How can audio deepfake detectors identify and localize manipulation when genuine and fake content coexist?
A whole-clip label can conceal which time intervals or overlapping sources provide evidence of manipulation. This makes mixed-authenticity audio decisions difficult to interpret and verify.
How can audio enhancement handle coupled real-world distortions while producing personalized, executable workflows?
Real-world recordings can contain interacting distortions, so correcting one artifact may affect others. Enhancement must also adapt to personalization requirements while producing workflows that are valid and executable.
How can audio-captioning datasets represent fine-grained acoustic detail and perceptual ambiguity for better audio retrieval?
Many audio-captioning datasets provide generic descriptions and only one caption per clip, even though listeners may describe the same sounds in different valid ways. Missing acoustic detail and semantic variation can limit models trained for audio retrieval and related audio-language tasks.
How can auscultation waveforms detect arteriovenous fistula dysfunction robustly across patients on resource-constrained devices?
Arteriovenous fistula dysfunction must be detected from sound recordings despite patient-specific variation and limited device compute. Conventional feature extraction may not transfer reliably across patients or after dimensionality reduction.
How can creators generate reusable multi-speaker voices and expressive audio scenes from instructions or reference recordings?
Audio production often requires coordinating several speakers, vocal styles, environments, and effects while preserving control over the resulting voices. Those voices may also need to be reused across creative projects.
How can diffusion-based music generators steer pitch content without retraining or modifying the base model?
Diffusion music generators provide limited direct control over the notes or pitch sequences in their outputs. The difficulty is imposing a desired pitch structure while preserving the generator and avoiding costly retraining.
How can discrete audio tokenizers preserve semantics and acoustic fidelity for both understanding and generation?
Audio tokenizers must compress continuous sound into manageable discrete sequences without losing either high-level content or acoustic detail. Separate semantic and acoustic streams can also introduce redundancy or misalignment when the same representation must support analysis and synthesis.
How can environmental-audio generators provide semantic control under limited compute and training data?
Generating everyday sounds with meaningful semantic control can require substantial computation and large, carefully curated datasets. These requirements make controllable environmental-audio synthesis difficult when training resources and data are limited.
How can multimodal models infer directorial intent from audiovisual choices rather than merely recognize events?
Recognizing what happens in a film does not explain why its lighting, composition, editing, dialogue, music, or sound were chosen. The central difficulty is connecting these audiovisual signals to the communicative meaning of filmmaking decisions.
How can multiple-choice music audio-language models estimate uncertainty well enough to abstain without costly ensembles or retraining?
Multiple-choice evaluation forces a model to choose even when its musical understanding is weak, so a lucky guess can resemble genuine competence. A single predictive distribution provides limited evidence about answer reliability, while independently trained ensembles are expensive.
How can tangible interaction communicate diffusion-model denoising without conventional technical explanation?
The iterative denoising process is difficult to encounter when audiences see only finished outputs, while conventional explanations may not support embodied exploration.
How can watermarks in generated speech remain detectable after neural audio transformations?
Watermarks embedded in speech can become unreliable after neural codecs, denoising, vocoding, and other processing used in storage or transmission. This undermines efforts to verify audio provenance while preserving an imperceptible listening experience.
How can we analyze closed digital audio processors when their internal signal-processing steps are inaccessible?
Closed-architecture processors expose output characteristics while concealing the internal operations and interacting systems that produce them. This opacity makes it difficult to connect implementation history with analytical understanding and conscious signal-processing design.
How can we determine whether large audio-language models exhibit human-like responses to auditory illusions?
Auditory illusions expose perceptual biases, but benchmarks for large audio-language models have largely emphasized visual illusions or general audio tasks. This makes it difficult to determine whether model responses reflect human-like perception or fidelity to the acoustic signal.