Home
Topics
Search
Library
Get Started
Research papers, read to you in five minutes.
Research papers, read to you in five minutes.
Every new AI paper that matters, distilled into a short audio episode.
Every new AI paper that matters, distilled into a short audio episode. Follow topics, listen on your commute, skim the recap when you're back.
Today's lead
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
Now playing · Today's lead
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
This week
all topics →
All
Audio/Speech
RAG
Agents
Audio Processing
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
Agents · Aug 26
Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence
Agents · Aug 21
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
Inference Optimization · Aug 21
Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs
Multimodal · Aug 20
4DAnyone: Create Anyone in 4D from a Casual Monocular Video
Multimodal · Aug 20
Never fall behind the literature again.
Sign up free
Agents · Aug 26
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
Voice agents can't afford 2-second memory lookups before replying, so this splits memory into a routing index over facts plus a separate persona/affect graph, matching schemas against partial transcripts during VAD silence. Hits 91.2 on LoCoMo with 430 tokens in 134ms, versus baselines needing 1,899 tokens.
Agents · Aug 21
Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence
When your single ReAct loop stops scaling, the instinct is a smarter agent; this survey argues the bottleneck is organization, not intelligence. It reframes multi-agent design as Graph Engineering: making task DAGs, agent capabilities, and runtime state explicit objects the runtime schedules, checkpoints, and rolls back on.
Inference Optimization · Aug 21
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
TLive-Omni tackles livestream commerce assistants that must ground answers in audio, video, and overlays at specific moments under latency pressure. Its twist: interleave audio-video tokens per time-grid, and use GRPO with a format reward that actively suppresses visible chain-of-thought instead of rewarding it.
Multimodal · Aug 20
Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs
OraRL fixes a subtle bug in GRPO post-training: when you inject a ground-truth annotation as a bonus rollout, it poisons the mean baseline and flips 22% of positive advantages negative. Keeping the oracle out of the baseline drops inversions to 0.3% and beats GRPO by 2+ points across tasks.
Multimodal · Aug 20
4DAnyone: Create Anyone in 4D from a Casual Monocular Video
Turning one phone video into a free-viewpoint avatar breaks because diffusion models can't hold 16+ novel views in one attention pass. 4DAnyone compresses accumulated reference views into a fixed token budget and rotates target-view groupings during high-noise steps, letting global structure propagate even when memory can't.
Never fall behind the literature again.
Free account. Follow topics, build your queue.
Sign up free
Agents · Aug 26
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
Voice agents can't afford 2-second memory lookups before replying, so this splits memory into a routing index over facts plus a separate persona/affect graph, matching schemas against partial transcripts during VAD silence. Hits 91.2 on LoCoMo with 430 tokens in 134ms, versus baselines needing 1,899 tokens.
Agents · Aug 21
Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence
When your single ReAct loop stops scaling, the instinct is a smarter agent; this survey argues the bottleneck is organization, not intelligence. It reframes multi-agent design as Graph Engineering: making task DAGs, agent capabilities, and runtime state explicit objects the runtime schedules, checkpoints, and rolls back on.
Inference Optimization · Aug 21
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
TLive-Omni tackles livestream commerce assistants that must ground answers in audio, video, and overlays at specific moments under latency pressure. Its twist: interleave audio-video tokens per time-grid, and use GRPO with a format reward that actively suppresses visible chain-of-thought instead of rewarding it.
Multimodal · Aug 20
Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs
OraRL fixes a subtle bug in GRPO post-training: when you inject a ground-truth annotation as a bonus rollout, it poisons the mean baseline and flips 22% of positive advantages negative. Keeping the oracle out of the baseline drops inversions to 0.3% and beats GRPO by 2+ points across tasks.
Multimodal · Aug 20
4DAnyone: Create Anyone in 4D from a Casual Monocular Video
Turning one phone video into a free-viewpoint avatar breaks because diffusion models can't hold 16+ novel views in one attention pass. 4DAnyone compresses accumulated reference views into a fixed token budget and rotates target-view groupings during high-noise steps, letting global structure propagate even when memory can't.
Never fall behind the literature again.
Free account. Follow topics, build your queue.
Sign up free
Browse topics
Audio/Speech 3
RAG 15
Agents 54
Audio Processing 2
Inference Optimization 33
Information Retrieval 1
All topics →
© 2026 r*cap — research papers as 5-minute podcasts
Terms
Privacy