Get Started
Topic · 3 recaps

Audio & Speech Processing

Speech recognition, synthesis, enhancement, and audio signal processing — the engineering counterpart to the Audio/Speech theme.
PostsQuestions
Home
Topics
Search
Library
Questions researchers are working onFollow a question through Rcap’s explanations and the latest papers addressing it.
How can acoustic sensing recover each person’s 3D pose despite overlapping motion signatures and inter-person reflections?Motion changes from several people are superimposed in the acoustic signal, while reflections create propagation delays that blur which changes belong to whom and when.How can an accessible humanoid robot integrate multimodal AI for real-world human interaction?Real-world human-robot interaction requires coordinating visual, gestural, and spoken inputs with physical manipulation. Integrating these capabilities on an accessible humanoid platform also requires maintaining accurate, timely control across the system.How can an autonomous audio system evolve sonic behavior without external data or post-initialization supervision?Without external data or feedback after initialization, an audio system must both produce changing sound and determine how its internal parameters should change. The difficulty is sustaining autonomous sonic evolution without simply becoming static or unstable.How can assistants remember and reason about how users sounded across long, multi-session conversations?Transcripts preserve words but can discard emotion labels, prosody descriptors, and voice events. Assistants working across long, multi-session histories may therefore fail on questions whose answers depend on how a user spoke.How can audio deepfake detectors identify and localize manipulation when genuine and fake content coexist?A whole-clip label can conceal which time intervals or overlapping sources provide evidence of manipulation. This makes mixed-authenticity audio decisions difficult to interpret and verify.How can audio enhancement handle coupled real-world distortions while producing personalized, executable workflows?Real-world recordings can contain interacting distortions, so correcting one artifact may affect others. Enhancement must also adapt to personalization requirements while producing workflows that are valid and executable.How can audio-captioning datasets represent fine-grained acoustic detail and perceptual ambiguity for better audio retrieval?Many audio-captioning datasets provide generic descriptions and only one caption per clip, even though listeners may describe the same sounds in different valid ways. Missing acoustic detail and semantic variation can limit models trained for audio retrieval and related audio-language tasks.How can audio-video diffusion models preserve intended conditioning when biased cross-modal attention reroutes semantics?In audio-video diffusion generation, cross-attention among text, audio, and video can route semantics bidirectionally rather than respecting intended conditioning. Learned biases may cause one modality to override prompts, producing visually canonical but semantically incorrect outputs.How can auscultation waveforms detect arteriovenous fistula dysfunction robustly across patients on resource-constrained devices?Arteriovenous fistula dysfunction must be detected from sound recordings despite patient-specific variation and limited device compute. Conventional feature extraction may not transfer reliably across patients or after dimensionality reduction.How can browser-based remote voice studies preserve audio provenance and integrity through capture, transfer, and transformation?Remote voice studies may retain a final audio file without reliable evidence of how it was captured, transferred, processed, or accepted. Missing provenance and integrity checks make it difficult to trace recording artifacts or detect transformations that could affect later analysis.How can causal spatial filters track moving speakers when only their initial directions are known?A spatial filter tuned to a speaker’s starting direction can lose the target as the speaker moves. Frame-wise causal processing also limits the information available for updating the direction estimate.How can co-speech gesture generation preserve semantic grounding and speech alignment without sacrificing biomechanical smoothness?Co-speech gesture systems must express lexical meaning while timing movements to speech. Semantic gestures and rhythmic beat gestures can compete, producing weak grounding, poor alignment, or jittery, physically implausible motion.How can conversational speech emotion recognition use cross-speaker context without confusing each speaker’s emotional trajectory?Emotion in dialogue depends on both a speaker’s own prior turns and what other participants say. Treating every adjacent utterance as one sequence can blur these distinct sources of temporal evidence across different time scales.How can creators generate reusable multi-speaker voices and expressive audio scenes from instructions or reference recordings?Audio production often requires coordinating several speakers, vocal styles, environments, and effects while preserving control over the resulting voices. Those voices may also need to be reused across creative projects.How can diffusion-based music generators steer pitch content without retraining or modifying the base model?Diffusion music generators provide limited direct control over the notes or pitch sequences in their outputs. The difficulty is imposing a desired pitch structure while preserving the generator and avoiding costly retraining.How can discrete audio tokenizers preserve semantics and acoustic fidelity for both understanding and generation?Audio tokenizers must compress continuous sound into manageable discrete sequences without losing either high-level content or acoustic detail. Separate semantic and acoustic streams can also introduce redundancy or misalignment when the same representation must support analysis and synthesis.How can environmental-audio generators provide semantic control under limited compute and training data?Generating everyday sounds with meaningful semantic control can require substantial computation and large, carefully curated datasets. These requirements make controllable environmental-audio synthesis difficult when training resources and data are limited.How can fine-grained audio-visual segmentation learn new classes continually without semantic drift or co-occurrence confusion?Sequentially learning fine-grained audio-visual segmentation classes can cause sounding objects to be treated as background in later tasks. Frequent class co-occurrences can also make visually or acoustically related objects difficult to distinguish.How can full-duplex dialogue models learn natural acoustic turn-taking without degrading semantic responses?Full-duplex systems must decide when to listen, speak, or yield while generating semantically appropriate responses. Synthetic text supervision misses the fine-grained acoustic timing of human conversation, while changing turn-taking behavior can disrupt semantic capability.How can full-duplex voice agents infer role-implied behavior while managing overlapping speech and conflicting instructions in real time?A role or persona may imply when an agent should listen, backchannel, interrupt, take the floor, or yield without stating those rules explicitly. The agent must infer and enact these behaviors during overlapping speech while resolving directives that may conflict with one another or with safety requirements.
Previous
1 / 3
Next