Home
Topics
Search
Library
Get Started
Research papers, read to you in five minutes.
Research papers, read to you in five minutes.
The day's top AI papers, distilled. The standouts become short audio episodes.
The day's top AI papers, distilled into clear written recaps. The standouts become short audio episodes for your commute.
Today's lead
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
Now playing · Today's lead
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
This week
all topics →
All
Agents
Audio Processing
RAG
Audio/Speech
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
Agents · Aug 26
Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence
Agents · Aug 21
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
Inference Optimization · Aug 21
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
Agents · Aug 20
4DAnyone: Create Anyone in 4D from a Casual Monocular Video
Multimodal · Aug 20
Never fall behind the literature again.
Sign up free
Agents · Aug 26
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
Voice agents can't afford 2-second memory lookups before replying, so this splits memory into a routing index over facts plus a separate persona/affect graph, matching schemas against partial transcripts during VAD silence. Hits 91.2 on LoCoMo with 430 tokens in 134ms, versus baselines needing 1,899 tokens.
Agents · Aug 21
Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence
When your single ReAct loop stops scaling, the instinct is a smarter agent; this survey argues the bottleneck is organization, not intelligence. It reframes multi-agent design as Graph Engineering: making task DAGs, agent capabilities, and runtime state explicit objects the runtime schedules, checkpoints, and rolls back on.
Inference Optimization · Aug 21
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
TLive-Omni tackles livestream commerce assistants that must ground answers in audio, video, and overlays at specific moments under latency pressure. Its twist: interleave audio-video tokens per time-grid, and use GRPO with a format reward that actively suppresses visible chain-of-thought instead of rewarding it.
Agents · Aug 20
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
SWE-bench Science tests coding agents on 119 real scientific repo repairs where private reverse-checks catch hard-coded fixes; no agent clears 50% pass@1. The kicker: adding domain context lifted a weak model 7 points but dropped a stronger one 4, anchoring it on plausible-wrong explanations.
Multimodal · Aug 20
4DAnyone: Create Anyone in 4D from a Casual Monocular Video
Turning one phone video into a free-viewpoint avatar breaks because diffusion models can't hold 16+ novel views in one attention pass. 4DAnyone compresses accumulated reference views into a fixed token budget and rotates target-view groupings during high-noise steps, letting global structure propagate even when memory can't.
Never fall behind the literature again.
Free account. Follow topics, build your queue.
Sign up free
Agents · Aug 26
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
Voice agents can't afford 2-second memory lookups before replying, so this splits memory into a routing index over facts plus a separate persona/affect graph, matching schemas against partial transcripts during VAD silence. Hits 91.2 on LoCoMo with 430 tokens in 134ms, versus baselines needing 1,899 tokens.
Agents · Aug 21
Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence
When your single ReAct loop stops scaling, the instinct is a smarter agent; this survey argues the bottleneck is organization, not intelligence. It reframes multi-agent design as Graph Engineering: making task DAGs, agent capabilities, and runtime state explicit objects the runtime schedules, checkpoints, and rolls back on.
Inference Optimization · Aug 21
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
TLive-Omni tackles livestream commerce assistants that must ground answers in audio, video, and overlays at specific moments under latency pressure. Its twist: interleave audio-video tokens per time-grid, and use GRPO with a format reward that actively suppresses visible chain-of-thought instead of rewarding it.
Agents · Aug 20
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
SWE-bench Science tests coding agents on 119 real scientific repo repairs where private reverse-checks catch hard-coded fixes; no agent clears 50% pass@1. The kicker: adding domain context lifted a weak model 7 points but dropped a stronger one 4, anchoring it on plausible-wrong explanations.
Multimodal · Aug 20
4DAnyone: Create Anyone in 4D from a Casual Monocular Video
Turning one phone video into a free-viewpoint avatar breaks because diffusion models can't hold 16+ novel views in one attention pass. 4DAnyone compresses accumulated reference views into a fixed token budget and rotates target-view groupings during high-noise steps, letting global structure propagate even when memory can't.
Never fall behind the literature again.
Free account. Follow topics, build your queue.
Sign up free
Browse topics
Agents 61
Audio Processing 2
RAG 16
Audio/Speech 3
Inference Optimization 37
Multimodal 41
All topics →
© 2026 r*cap — research papers as 5-minute podcasts
Terms
Privacy
r*cap - Featured on Startup FameFeatured on Twelve ToolsFeatured on tinyshelf