ReaLVR teaches a multimodal model’s hidden “thinking” tokens to actually encode the image evidence that supports the answer, by contrasting correct vs. wrong answers to decide where to supervise and relevant vs. mismatched image regions to decide what to preserve.
Multimodal models increasingly do “thinking” in continuous Latent Visual Reasoning (LVR) rather than written-out reasoning steps. The idea is that for spatial or fine-grained visual problems, dense latent vectors can carry more information than a few sentences of English.
The practical problem the authors flag: nobody actually checks whether those latent vectors are reading the image. Standard training uses Group Relative Policy Optimization (GRPO), which only rewards the final text answer. So the model gets no direct signal about which latent positions should carry visual evidence, or what evidence they should carry. The authors call this the latent evidence-credit gap.
Their diagnostic makes this concrete. They edit images in ways that should flip the answer (change a color, remove an object, swap a shape, swap left/right). The ground-truth answer flips on roughly 81\u201386% of pairs, but a vanilla Latent Visual Reasoning (LVR) model changes its prediction on only 5.7\u201313.1%. The latent trajectory barely moves when the evidence that matters changes. So the latent tokens are mostly inheriting scene gist, not tracking the decisive cue.
ReaLVR adds a second training objective alongside the usual answer reward. The high-level picture:
1.
Regenerate the latent trajectory with the current model, keeping gradients through the recurrence. This matters because at inference the model generates its own latent tokens, so supervision has to flow through that same generation process. This is an On-Policy Distillation move.
2.
Decide what each latent token should encode (visual contrast). Pool image-token features inside the annotated region of interest to make a positive visual prototype. Pool features from mismatched examples as negatives. Push each latent state’s cosine similarity toward the positive and away from the hardest negative. If no region is annotated, use the whole image.
3.
Decide which latent positions deserve the strongest push (answer contrast). Teacher-force the correct answer after the latent span and record how much each answer token attends to each latent position. Do the same for the model’s own sampled wrong answers. Subtract: positions that the correct answer reads more than wrong answers do get a bigger weight. This cancels out attention that’s just about formatting or sentence-start behavior.
4.
Detach the weights when computing the visual loss. Otherwise the model finds a shortcut: it can lower the loss by routing weight away from hard positions instead of actually learning the evidence. The Stop-gradient blocks that path.
5.
Architecture and inference are unchanged. The whole mechanism is a training-time add-on.
# Per training example, alongside standard GRPO
z = regenerate_latents(model, x) # keep grads through recurrence
p_pos = pool(image_tokens, roi_mask) # what to preserve
N_neg = pool_from_mismatched_examples(batch) # visual negatives
r_pos = attention_from(correct_answer, z) # teacher-force correct
r_neg = mean(attention_from(y, z) for y in wrong_answers_sampled)
gamma = relu(r_pos - r_neg) # where to supervise
w = eta/K + (1-eta) * gamma # mix with uniform floor
for t in range(K):
margin = cos(z[t], p_pos) - max(cos(z[t], p) for p in N_neg)
loss += stop_grad(w[t]) * relu(m_ev - margin)
Headline accuracy. On Qwen2.5-VL-7B, ReaLVR’s five-benchmark average is 63.7%, versus 60.4% for an Latent Visual Reasoning (LVR)-RL baseline that uses the same data and compute, and 62.9% for the strongest competing latent-reasoning entry, ILVR. The suite covers MMVP, BLINK, HRBench at 4K and 8K, and MME-RealWorld. The largest single-task jump over an Latent Visual Reasoning (LVR)-SFT baseline is +8.4 points on MMVP.
Scaling. The same recipe applies without change to larger backbones. On Qwen3-VL-30B ReaLVR reaches 65.2% average (+1.1 over LVR-RL). On Qwen3-VL-235B (evaluated on only three of the five tasks), ReaLVR improves over LVR-SFT on each. The authors frame this as the first demonstration of continuous latent visual reasoning trained at frontier scale, though the 235B evaluation is partial and compute is substantial (800 MI250X accelerators for the 235B run).
Cross-family. On InternVL3-8B and Gemma 3-12B, ReaLVR beats LVR-RL by 3.1 and 2.6 points respectively.
Mechanism checks are the more interesting evidence. Three results support the claim that the latents are now doing work, rather than just that the number went up:
•
On the image-edit test, ReaLVR flips its answer where LVR does not; mean latent distance between original and edited inputs rises to 0.13\u20130.34, versus below 0.0005 for LVR. The paper does not report ReaLVR’s own correct-flip rate aggregated against the 81\u201386% ground-truth rate in the main text, so this is directional rather than a matched comparison.
•
Target-region attention enrichment reaches ~2x on middle layers for ReaLVR (vs ~1.3 for LVR), meaning the answer tokens attend roughly twice as much to the annotated region as to equal-area background. Caveat: correct and incorrect answers show similar enrichment, so looking at the right place is necessary but not sufficient.
•
A fixed-context replacement test substitutes the 8 most answer-attended latent tokens; ReaLVR’s correct-answer probability drops from 0.70 to 0.59, a larger drop than LVR or Monet. The latents are load-bearing locally.
Ablations (Appendix F). Every piece matters. Removing visual negatives costs 1.6 points. Training on the saved rollout latents instead of regenerating costs 1.8 points. Removing the stop-gradient on weights costs 1.1 points, exactly the shortcut the authors predicted. Uniform routing (no answer contrast) costs 1.3 points.
If you are training a multimodal model that uses continuous latent tokens between the image and the answer: the paper is a strong argument that final-answer rewards alone are insufficient to make those tokens track visual evidence. The diagnostic (compare prediction-flip rate under answer-changing image edits to ground-truth flip rate) is cheap to run and would tell you quickly whether your own latent tokens are responsive or inert.
If you are considering adding ReaLVR-style supervision: the prerequisites are nontrivial. You need region-of-interest annotations (or the whole-image fallback, which is weaker), access to the vision tower’s token features for pooling prototypes, and the ability to teacher-force candidate answers after a regenerated latent span to extract attention. This is a research-training setup, not a hosted-API intervention.
If you don’t do latent reasoning: the useful transferable idea is the diagnostic design, not the loss. Rank intermediate representations by how much the answer attends to them, then intervene on the top-k while holding the rest fixed. If the answer probability doesn’t move much, your “reasoning” tokens are decorative. This logic applies to any pipeline with intermediate representations.
One thing worth testing before committing: the latent budget sweep shows the best K is task-dependent (K=4 for MME-RealWorld, K=16 for BLINK). A fixed budget leaves accuracy on the table, so an adaptive-length scheme is a natural next experiment the paper does not run.
•
The target-region alignment is similar for correct and incorrect answers, so attending to the right region does not predict correctness. Spatial grounding is a prerequisite the method improves, not a sufficient condition for better answers.
•
When no region-of-interest annotation is available, the positive prototype is the whole image, which the authors acknowledge is less discriminative. Benefits may shrink on datasets without region labels.
•
Gains over the strongest latent-reasoning baseline on the 7B five-task average are modest (+0.8 over ILVR). The bigger case for ReaLVR is the mechanism diagnostics, not raw leaderboard delta.
•
The 235B result rests on three benchmarks, not the full five. “First at frontier scale” is a scale claim, not a comprehensive-evaluation claim.
•
Training compute is large (800 MI250X accelerators for 235B, 256 for smaller models). Reproducing the frontier-scale result outside a well-resourced lab is not realistic.
•
Inference uses a fixed latent budget. The ablation shows per-task optima differ, so a single deployed K is a compromise.