Home
Topics
Search
Library
Get Started
Research papers, read to you in five minutes.
Research papers, read to you in five minutes.
Every new AI paper that matters, distilled into a short audio episode.
Every new AI paper that matters, distilled into a short audio episode. Follow topics, listen on your commute, skim the recap when you're back.
Today's lead
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
Now playing · Today's lead
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
This week
all topics →
All
LLM Training
Information Retrieval
Reasoning
Multimodal
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
Inference Optimization · Aug 21
Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence
Agents · Aug 21
4DAnyone: Create Anyone in 4D from a Casual Monocular Video
Multimodal · Aug 20
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization
Reasoning · Aug 17
Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision
Image Generation · Aug 17
Never fall behind the literature again.
Sign up free
Inference Optimization · Aug 21
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
TLive-Omni tackles livestream commerce assistants that must ground answers in audio, video, and overlays at specific moments under latency pressure. Its twist: interleave audio-video tokens per time-grid, and use GRPO with a format reward that actively suppresses visible chain-of-thought instead of rewarding it.
Agents · Aug 21
Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence
When your single ReAct loop stops scaling, the instinct is a smarter agent; this survey argues the bottleneck is organization, not intelligence. It reframes multi-agent design as Graph Engineering: making task DAGs, agent capabilities, and runtime state explicit objects the runtime schedules, checkpoints, and rolls back on.
Multimodal · Aug 20
4DAnyone: Create Anyone in 4D from a Casual Monocular Video
Turning one phone video into a free-viewpoint avatar breaks because diffusion models can't hold 16+ novel views in one attention pass. 4DAnyone compresses accumulated reference views into a fixed token budget and rotates target-view groupings during high-noise steps, letting global structure propagate even when memory can't.
Reasoning · Aug 17
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization
Multi-reward RLVR wastes gradient reinforcing objectives the model already maxed out. SA-MRPO reweights each objective's advantage by (1 − saturation)^γ before normalization, redirecting pressure to unsolved rewards — +9.2pp on AMC23 when a length reward saturates, with γ≈0.5 as the sweet spot.
Image Generation · Aug 17
Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision
Instruction-tuned image editors scale the wrong axis: more source images under 10-25 coarse operation buckets. ConceptEdit flips it — balance across 1,028 fine-grained edit concepts and pack spatially disjoint edits per target, hitting target quality with ~1.5× fewer samples and lifting single-edit categories too.
Never fall behind the literature again.
Free account. Follow topics, build your queue.
Sign up free
Inference Optimization · Aug 21
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
TLive-Omni tackles livestream commerce assistants that must ground answers in audio, video, and overlays at specific moments under latency pressure. Its twist: interleave audio-video tokens per time-grid, and use GRPO with a format reward that actively suppresses visible chain-of-thought instead of rewarding it.
Agents · Aug 21
Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence
When your single ReAct loop stops scaling, the instinct is a smarter agent; this survey argues the bottleneck is organization, not intelligence. It reframes multi-agent design as Graph Engineering: making task DAGs, agent capabilities, and runtime state explicit objects the runtime schedules, checkpoints, and rolls back on.
Multimodal · Aug 20
4DAnyone: Create Anyone in 4D from a Casual Monocular Video
Turning one phone video into a free-viewpoint avatar breaks because diffusion models can't hold 16+ novel views in one attention pass. 4DAnyone compresses accumulated reference views into a fixed token budget and rotates target-view groupings during high-noise steps, letting global structure propagate even when memory can't.
Reasoning · Aug 17
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization
Multi-reward RLVR wastes gradient reinforcing objectives the model already maxed out. SA-MRPO reweights each objective's advantage by (1 − saturation)^γ before normalization, redirecting pressure to unsolved rewards — +9.2pp on AMC23 when a length reward saturates, with γ≈0.5 as the sweet spot.
Image Generation · Aug 17
Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision
Instruction-tuned image editors scale the wrong axis: more source images under 10-25 coarse operation buckets. ConceptEdit flips it — balance across 1,028 fine-grained edit concepts and pack spatially disjoint edits per target, hitting target quality with ~1.5× fewer samples and lifting single-edit categories too.
Never fall behind the literature again.
Free account. Follow topics, build your queue.
Sign up free
Browse topics
LLM Training 35
Information Retrieval 1
Reasoning 25
Multimodal 36
Agents 53
Inference Optimization 33
All topics →
© 2026 r*cap — research papers as 5-minute podcasts
Terms
Privacy