Get Started
Home
Topics
Search
Library
Research questionHow can speaker-attributed ASR identify who said what as speech arrives with low latency?Streaming speaker-attributed ASR must produce both the transcript and the speaker identity from partial audio. Delaying decisions until a recording is complete conflicts with the low-latency needs of interactive systems.
AI
Audio & Speech
Audio & Speech Processing
Evaluation & Benchmarks
Research Paper
Technology
Latest papersRecent research connected to this question, newest first.VibeVoice-ASR-Streaming Technical ReportThe source concerns LLM-based end-to-end streaming speaker-attributed ASR using fixed-size audio chunks, limited lookahead, and prior text, without a separate diarization stage. It reports results for 1.5B and 7B models across multiple transcription and speaker-attribution evaluation settings, but does not specify deployment latency measurements or access requirements.research paper · Sep 10, 2026
Related questions
How can zero-shot voice conversion transfer an unseen speaker’s identity while preserving content in low-latency streaming?How can assistants remember and reason about how users sounded across long, multi-session conversations?How can speech synthesis produce natural dubbing and full-duplex dialogue without forced alignment or explicit duration prediction?How can speech recognition avoid hallucinated transcripts on non-speech without degrading genuine speech?