Get Started
Home
Topics
Search
Library
6 min read · Agents · Inference Optimization · Added Oct 8 · Paper published Oct 1, 2026

Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents

Source: research paper via Hugging Face Daily Papers
Small multimodal agents waste most per-turn latency decoding formulaic reasoning like “I should search for this.” SSR replaces freeform chain-of-thought with parallel scoring over six hand-written rationales, cutting reasoning latency 90%+ on Qwen3-VL 2B/4B search agents while matching accuracy — and RL-trainable without an SFT warm-up.
TL;DR
SSR replaces an agent’s freeform chain-of-thought with picking one of six pre-written rationales by scoring each in parallel against the current context, cutting per-turn reasoning latency by over 90% on small multimodal search agents without hurting task success.
Why It Matters
You’re building a small (2B-4B parameter) multimodal agent that answers visual questions by calling tools: reverse image search, text search, image zoom. The standard recipe is ReAct: each turn, the model writes a few sentences of reasoning (“I don’t recognize this building, I should run image search first”), then emits a tool call. That reasoning text has to be decoded token by token, and for small models it tends to be long, repetitive, and not actually that informative. You pay real latency for prose that mostly says “I should search for X.”
Existing fixes either shrink the reasoning text (Chain-of-Draft, token budgets) or train the model to skip reasoning when it can. Both still generate autoregressively. The authors from Apple and CMU notice something simpler: across multimodal search questions, the high-level moves recur. “Entity unknown, look it up by image” covers a huge fraction of turns. If the space of useful rationales is small, why generate from the full language distribution every time?
How It Works
The core move: hand-write a tiny library of six reusable rationales (search by image, search by text, zoom into the image, answer from retrieved evidence, answer from the image, answer from general knowledge). At each turn, instead of generating reasoning, the model scores each candidate and picks one.
Scoring reuses the model itself. No classifier head, no separate retriever. For each candidate, compute the average log-probability its tokens would get under the current context, then softmax across candidates to get a selection distribution. Because the candidate text is fixed in advance, Teacher Forcing (chunk-wise causal) lets the model score every token in every candidate in a single parallel forward pass, all sharing the context’s KV cache. The selected candidate’s text is then pasted into the context as the “reasoning” that conditions the next action (the actual search query, crop coordinates, or final answer).
So autoregressive generation only happens for the action, not the reasoning. Pseudocode for one turn:
def ssr_turn(history, candidates, model): # score all candidates in parallel, sharing history KV cache scores = model.score_parallel(history, candidates) # length-normalized logprobs probs = softmax(scores / tau) chosen = sample(candidates, probs) # one categorical draw context = history + chosen # paste chosen text as reasoning action = model.generate(context) # autoregressive, only for action return chosen, action
Training works with either SFT or RL. For Group Relative Policy Optimization (GRPO) and its variants GSPO and SAPO, the token-level importance ratio over reasoning tokens collapses into a single categorical ratio over the selected candidate, while action tokens keep their usual per-token ratios. Unlike prior methods that need an SFT warm-up to teach the model to reason in a specific format, SSR can be RL-trained from scratch because the library already constrains the reasoning space.
What They Found
Using Qwen3-VL 2B and 4B backbones across seven multimodal search benchmarks (MMSearch, HR-MMSearch, FVQA-test, InfoSeek, SimpleVQA, LiveVQA, MAT-Search):
•
The 4B SSR model trained with GRPO hits 61.37% average success, matching the best 4B baseline TAPO at 61.25% and beating other 4B search agents the authors compare to.
•
Per-turn reasoning latency drops by over 90% vs the matched freeform-reasoning version of the same model. Per-question total model latency drops 28-54% depending on training objective.
•
At the 95th percentile, SSR’s reasoning time is 71 ms (10 ms above its mean). Baselines range from ~0.8s to ~1.9s at p95. Scoring a fixed library has predictable cost; generating variable-length prose does not.
•
Against four efficient-reasoning baselines (Chain-of-Draft, Sketch-of-Thought, efficiency-reward RL, Probe & Prefill), SSR is simultaneously the most accurate and roughly 7× faster on reasoning.
Two ablations worth noting. Index-only selection (scoring the candidate’s number [1] instead of its full text) loses 3.9 points of average success, so the semantics of the candidate tokens matter for picking the right one, not just the identity. Library size matters a lot at the low end: going from one generic candidate (“I am a helpful assistant”) to two (answer vs. tool call) adds 17.27 points; going from one to six adds about 24. Finally, the selected rationale genuinely drives behavior: in 99.9% of turns the generated action type matches the type the chosen candidate implies, and the mix of chosen candidates shifts sensibly across benchmarks (image-centric benchmarks pick “examine image detail” more; simple VQA picks “answer from image” more).
One honest wrinkle: SSR beats freeform GRPO at both sizes but shows modest regressions under GSPO and SAPO. The authors attribute this to their selector-gradient approximation (they hold competing candidates’ scores fixed at rollout values to avoid backprop through all candidates), which drifts across multiple updates per rollout.
What’s Useful
•
If you’re shipping a small on-device multimodal agent and your traces show lots of short, formulaic reasoning blocks before each tool call, this is a strong efficiency pattern to try. The win comes from eliminating autoregressive decoding of reasoning, so it’s largest when reasoning tokens are a meaningful fraction of your per-turn cost.
•
The approach needs logprob access to the model, including the ability to score pre-specified token sequences via teacher-forced prefill. The paper uses SGLang for this. If you only have hosted chat-completion access without logprob scoring over arbitrary continuations, you cannot run SSR as described.
•
The library is manually written. The authors suggest sampling successful freeform traces and clustering recurring rationales to build one; worth testing on your own domain before assuming six generic entries transfer. The six used here are tuned to search-agent action types.
•
For RL training specifically, SSR removes the usual need for an SFT warm-up to lock the model into a reasoning format. That’s a meaningful simplification if you were previously running two-stage pipelines.
•
Project page: ssr-webpage. The paper does not state a code or weights release in the supplied text.
Caveats
•
Results are on 2B and 4B Qwen3-VL backbones in a search-agent setting with three tools. The library is tailored to this action space; a general-purpose coding or planning agent would need a different (and likely larger) candidate set, and the paper does not evaluate that.
•
Library coverage caps performance. If the right rationale for a turn isn’t in the library, the model can’t invent one. The size-ablation curve was still rising at six entries; the ceiling for this task family isn’t established.
•
The success-rate parity claim is “competitive with same-scale search agents,” not a universal SOTA. On HR-MMSearch and MAT-Search, freeform reasoning still beats SSR.
•
The selector-gradient approximation used for GSPO/SAPO loses fidelity across repeated updates on the same rollout batch, which the authors flag as the likely cause of their small regressions under those objectives.
Topics
Agents
Inference Optimization
Multimodal
Agents
Inference Optimization
Multimodal
Up next in Agents
From Evidence to Action: How Tool-Using Agents Fail
MiniCorp: The Last Mile of the AI Agent Firm
Don't miss new content
Log in to follow topics and personalize your feed.
Related topics you might like
Agents236 episodes
Multimodal128 episodes
Inference Optimization135 episodes