Get Started
Topic · 8 recaps

Audio & Speech

Speech recognition, text-to-speech, voice cloning, music generation, and audio understanding — across both transformer and diffusion architectures.
PostsQuestions
Home
Topics
Search
Library
Questions researchers are working onFollow a question through Rcap’s explanations and the latest papers addressing it.
How can an accessible humanoid robot integrate multimodal AI for real-world human interaction?Real-world human-robot interaction requires coordinating visual, gestural, and spoken inputs with physical manipulation. Integrating these capabilities on an accessible humanoid platform also requires maintaining accurate, timely control across the system.How can assistants remember and reason about how users sounded across long, multi-session conversations?Transcripts preserve words but can discard emotion labels, prosody descriptors, and voice events. Assistants working across long, multi-session histories may therefore fail on questions whose answers depend on how a user spoke.How can audio deepfake detectors identify and localize manipulation when genuine and fake content coexist?A whole-clip label can conceal which time intervals or overlapping sources provide evidence of manipulation. This makes mixed-authenticity audio decisions difficult to interpret and verify.How can audio enhancement handle coupled real-world distortions while producing personalized, executable workflows?Real-world recordings can contain interacting distortions, so correcting one artifact may affect others. Enhancement must also adapt to personalization requirements while producing workflows that are valid and executable.How can audio-captioning datasets represent fine-grained acoustic detail and perceptual ambiguity for better audio retrieval?Many audio-captioning datasets provide generic descriptions and only one caption per clip, even though listeners may describe the same sounds in different valid ways. Missing acoustic detail and semantic variation can limit models trained for audio retrieval and related audio-language tasks.How can browser-based remote voice studies preserve audio provenance and integrity through capture, transfer, and transformation?Remote voice studies may retain a final audio file without reliable evidence of how it was captured, transferred, processed, or accepted. Missing provenance and integrity checks make it difficult to trace recording artifacts or detect transformations that could affect later analysis.How can causal spatial filters track moving speakers when only their initial directions are known?A spatial filter tuned to a speaker’s starting direction can lose the target as the speaker moves. Frame-wise causal processing also limits the information available for updating the direction estimate.How can co-speech gesture generation preserve semantic grounding and speech alignment without sacrificing biomechanical smoothness?Co-speech gesture systems must express lexical meaning while timing movements to speech. Semantic gestures and rhythmic beat gestures can compete, producing weak grounding, poor alignment, or jittery, physically implausible motion.How can conversational speech emotion recognition use cross-speaker context without confusing each speaker’s emotional trajectory?Emotion in dialogue depends on both a speaker’s own prior turns and what other participants say. Treating every adjacent utterance as one sequence can blur these distinct sources of temporal evidence across different time scales.How can creators generate reusable multi-speaker voices and expressive audio scenes from instructions or reference recordings?Audio production often requires coordinating several speakers, vocal styles, environments, and effects while preserving control over the resulting voices. Those voices may also need to be reused across creative projects.How can discrete audio tokenizers preserve semantics and acoustic fidelity for both understanding and generation?Audio tokenizers must compress continuous sound into manageable discrete sequences without losing either high-level content or acoustic detail. Separate semantic and acoustic streams can also introduce redundancy or misalignment when the same representation must support analysis and synthesis.How can environmental-audio generators provide semantic control under limited compute and training data?Generating everyday sounds with meaningful semantic control can require substantial computation and large, carefully curated datasets. These requirements make controllable environmental-audio synthesis difficult when training resources and data are limited.How can full-duplex dialogue models learn natural acoustic turn-taking without degrading semantic responses?Full-duplex systems must decide when to listen, speak, or yield while generating semantically appropriate responses. Synthetic text supervision misses the fine-grained acoustic timing of human conversation, while changing turn-taking behavior can disrupt semantic capability.How can full-duplex voice agents infer role-implied behavior while managing overlapping speech and conflicting instructions in real time?A role or persona may imply when an agent should listen, backchannel, interrupt, take the floor, or yield without stating those rules explicitly. The agent must infer and enact these behaviors during overlapping speech while resolving directives that may conflict with one another or with safety requirements.How can joint audio-video generators follow script-specified timing for shot transitions and dialogue?Audio and video can remain synchronized while both occur at the wrong times relative to the script. This disrupts narrative structure when shot changes or dialogue are tied to specified moments.How can low-resource Thai TTS learn a fixed voice from synthetic speech while preserving pronunciation and prosody?With little speaker-specific data, a fixed-voice Thai synthesizer can avoid the inference cost of voice cloning, but synthetic targets may reproduce teacher errors or omit difficult text. Thai word boundaries, lexical tones, names, numbers, and code-switching make pronunciation and prosody particularly sensitive to these choices.How can multimodal models rely on images or audio rather than language shortcuts?A model can choose a plausible answer from wording alone while overlooking the image or audio. The challenge is making its answer depend on what it actually sees or hears.How can multimodal speech and language cues detect loneliness in older adults during telephone interviews?Loneliness may be expressed through both what older adults say and how they sound, but these cues are difficult to interpret consistently in conversational speech. Telephone-based analysis must connect linguistic and acoustic patterns with loneliness without equating association with diagnosis.How can multiple-choice music audio-language models estimate uncertainty well enough to abstain without costly ensembles or retraining?Multiple-choice evaluation forces a model to choose even when its musical understanding is weak, so a lucky guess can resemble genuine competence. A single predictive distribution provides limited evidence about answer reliability, while independently trained ensembles are expensive.How can non-invasive speech BCIs classify words for new users with only minutes of subject-specific MEG data?Speech BCIs can decode words more accurately when they have extensive recordings from the same person. Practical systems instead need to transfer across users and adapt from a clinically feasible amount of new-user data.
Previous
1 / 3
Next