OneStreamer keeps a short visual window plus text notes it writes to itself ([Proactive Hierarchical Caption Memory](glossary://phcm)), and learns when to speak vs. wait by only training on informative state-token moments ([Proactive State Transition Learning](glossary://pstl)), so a 4B streaming video model can answer about things that already scrolled off screen without hoarding old frames.
Say you’re building a wearable assistant or a live-stream monitor. Video comes in forever, but the model only has room for maybe the last 16 to 32 frames of visual tokens. The user might ask at minute 2 about something that happened at second 7, long after those pixels got evicted. You don’t know in advance what will matter.
Two common fixes both hurt. One: keep more frames (or compress them into visual memory tokens). That inflates context and, as the paper notes citing prior work, extensive visual history actually degrades perception of what’s happening right now. Two: use only a recent window and lose the past entirely. There’s also a separate timing problem: in a streaming model, most output tokens are “keep quiet” markers, so if you train naively, the model learns to shut up and rarely emits a response.
The closest named prior work the paper builds on for the response-timing mechanism is Streamo, which uses generative state tokens to control when to speak. OneStreamer accepts that framing and attacks both the memory-vs-perception tradeoff and the silence-token imbalance in one shared setup.
The idea is to make the model write itself notes as the video plays, in the same generation loop it uses to answer users. Instead of storing old frames, it stores short text records tagged with timestamps. Later, when a question arrives, those notes sit in context alongside the recent frames.
There are two kinds of notes, each triggered by a dedicated control token. </Observe> starts a dense local caption describing what was just seen (objects, actions, state changes). </Summary> starts a coarser recap of a completed event. For live Q&A, three more tokens drive timing: </Silence> (keep watching), </Standby> (relevant evidence is forming), and </Response> (now answer). All of this is one causal sequence of visual tokens, text, and control tokens. That bundle is Proactive Hierarchical Caption Memory.
At inference, the recipe is roughly:
for clip in stream:
visual_window.push(clip) # FIFO, keep last N frames
ctx = interleave(visual_window, text_history)
tok = model.predict_control(ctx)
if tok in (OBSERVE, SUMMARY):
note = model.generate_text(ctx)
text_history.append((timestamp, tok, note))
elif tok == RESPONSE and user_question_pending:
answer = model.generate_text(ctx)
emit(answer)
# else SILENCE / STANDBY: do nothing visible
Old frames fall out of the window, but the generated notes stay in text_history, so distant events survive as cheap text.
The second piece, Proactive State Transition Learning, is about training those control tokens. In a long sequence, </Silence> dominates. If you cross-entropy-train every state token equally, the model overfits to staying quiet. PSTL instead: (1) always supervise every token that initiates an output (every </Observe>, </Summary>, </Response>), and (2) for each kind of state transition in a sequence, subsample down to a shared quota so frequent transitions (mostly silence-to-silence) don’t swamp rare ones. Unselected state tokens stay in the sequence as context but are masked out of the state loss only. Caption and answer text keep full supervision throughout.
Training data comes from a synthesis pipeline the authors built: videos get multi-granularity captions from Gemini and Seed, verified for temporal grounding, then converted into streaming sequences where each caption is released only after its supporting frames. QA pairs get a calibrated response time: the earliest moment where a verifier model can answer correctly from the prefix. The combined corpus is OneStreamer-1M, roughly 1.16M records across captioning, proactive interaction, proactive QA, and perception/memory QA. The base model is Qwen3-VL-4B-Instruct, fine-tuned on 32 H200s.
The headline is on eight streaming video benchmarks (four for perception/memory: OVOBench, StreamingBench, OVBench, ODVBench; four for proactive response: ProactiveVQA, OmniMMI, OVO-Timing, ViSpeak). The 4B OneStreamer tops every one against models up to 11B, with the paper framing it as +25.0% average relative improvement over the Qwen3-VL baseline and +6.1% over the strongest competitor per benchmark.
The more informative results are the ablations:
•
PHCM beats both extremes, not just splits the difference. With the same recent 16-frame window, adding caption memory raises OVOBench Backward ASI from 63.5 to 71.6 vs. FIFO alone, and also beats the Full-visual-history variant on that metric (67.6). On real-time perception, PHCM (81.4) still edges FIFO (80.9) and clearly beats Full (70.9), which is the key “doesn’t hurt the present to remember the past” claim.
•
PSTL beats dense state supervision while supervising only 27.5% of state tokens. On the three proactive-response benchmarks reported, PSTL hits 48.7 / 36.6 / 41.6 vs. 26.1 / 30.8 / 1.5 for plain cross-entropy on all tokens. A random-sparse control at the same 27.5% ratio does much worse (26.6 / 27.4 / 1.4), so the gain is from which tokens are kept, not just sparsity. A focal-loss-on-all baseline from Streamo closes most of the gap but still trails on OmniMMI and OVO-Timing.
•
Latency. On a 360s clip, PHCM uses 4,308 context tokens and 0.124s time-to-first-token vs. 62,094 tokens and 4.56s for retaining full visual history. Over FIFO, the overhead for carrying caption memory is about 0.03s and 0.29 GB.
•
Caption supervision helps even without caption memory at inference. A variant trained without the 60K proactive caption examples loses on every setting, suggesting the streaming-caption targets act as fine-grained supervision for interpreting video prefixes, independent of their memory role.
Failure cases exist: in one, the model calls a printer being opened “closing” at 00:10 and never retracts within the shown window. Interpreting state changes and self-correcting remain weak spots.
•
If you’re building a streaming video assistant and already use a recent-frame window, the PHCM pattern (write timestamped text notes into your context as you go, rather than hoarding visual tokens) is a cheap, concrete design to try. The latency numbers suggest it’s close to free vs. a plain recent window and dramatically cheaper than full history. The caveat: you need to train the model to emit those </Observe> and </Summary> tokens in the right places. Prompting an off-the-shelf VLM to “describe what you see every few seconds” is not the same intervention. The ablation where a model without caption-data training is prompted the same way shows essentially no memory gain.
•
If you’re training any model that emits sparse decisions among dense “wait” tokens (streaming agents, interrupt detectors, trigger-word classifiers), the PSTL recipe is worth testing: keep every output anchor, cap each transition group by the largest output-anchor group’s size, mask the rest from the loss only. The random-sparse baseline failing at the same ratio is the evidence that this specific selection matters.
•
For dataset construction in streaming settings, the response-time calibration idea (find the earliest prefix where a verifier can answer correctly, then anchor the response there) is a reusable trick. The paper reports it moves response times earlier by 13.72s on average vs. end-of-interval labels, with 86.62% of records getting an earlier timestamp.
•
Deployment caveats: results are from a single 4B model fine-tuned on 32 H200s with a specific data mix. Transfer to other base models or smaller training budgets isn’t shown. All eight benchmarks are video-understanding evals, not real deployments, so “wins the benchmark suite” is not the same as “reliable in a wearable at 1 FPS over hours.” The paper doesn’t link a code or weights release in the supplied text; the project page is listed but artifact availability isn’t specified.
•
The main results compare a model trained on OneStreamer-1M against baselines trained on their own corpora. Some of the gain is data, some is method; the ablations isolate PHCM and PSTL but not the full data advantage against every competitor.
•
Caption memory is only as good as what the model writes. If </Observe> records are wrong (as in the printer failure case), that error is now persisted into the context and can mislead later answers. The paper doesn’t quantify how often this happens.
•
The 25% and 6.1% average-improvement framings are the authors’; interpret them as per-benchmark relative gains, not a single metric.
•
The PSTL supervision ratio (27.5%) is sequence-dependent since the quota is set per sequence from the largest output-anchor group. It’s not a tunable knob you set directly.