RL post-training of LLMs trades solution coverage for single-shot accuracy by pushing per-task success rates to all-pass or all-fail; the authors measure this with a Sharpening Tax metric and reduce it with a per-prompt adaptive-temperature sampler during RL.
You have a base LLM and its post-trained (RL-tuned, chat/instruct) sibling. Common wisdom says the post-trained one is strictly better for agent work: tool calls, multi-turn planning, following instructions. So you ship the post-trained version and give it one shot per task.
The paper’s concern: if you can afford many rollouts per task (parallel sampling, best-of-N with a verifier, scientific search), the post-trained model may actually solve fewer distinct tasks than the base model. Prior work saw this coverage loss on math and code but could be dismissed: base LLMs saw tons of math in pretraining, and final-answer grading rewards lucky-but-wrong reasoning. Agentic tasks (Berkeley Function Calling Leaderboard multi-turn tool calls, WebShop web navigation, ACEBench) are a cleaner test because intermediate environment states are checked, and tool-use skills are widely assumed to come from post-training.
The surprise: with a plain-text scaffold (a harness) that teaches the base model how to format tool calls, base models do agentic tasks competently and, given enough rollouts, cover more tasks than their post-trained versions.
The mechanism they identify is bimodalization. Think of each task having some success probability p under a policy. Base models spread tasks across intermediate p values, where extra rollouts keep converting failures to wins. Post-training pushes mass toward the extremes: tasks become either always-solved (p≈1) or always-failed (p≈0). Both extremes kill the value of retries. One always succeeds on try one, the other never succeeds.
To quantify this, they define Sharpening Tax. In plain English: measure the area between a model’s pass@K ceiling and its pass@k curve as k grows. That area is literally the expected number of failed attempts before the first success within budget K. If post-training shrinks this area versus the base model, retries help it less, and the tax is positive.
They report two variants: raw area, and a calibrated version normalized by the single-shot failure rate so it sits in [0,1].
Their fix during RL training is Posterior-Tempered Group Sampling. Idea: during training, track each prompt’s recent success rate with a Beta posterior, sample a difficulty estimate per prompt via Thompson sampling, then set that prompt’s rollout temperature. Hard prompts get heated (more exploration), easy prompts get cooled (more exploitation).
# per prompt x during RL training
s_x, f_x = discounted_success_fail_counts[x]
a = 2*p_target + s_x
b = 2*(1-p_target) + f_x
p_hat = Beta(a, b).sample() # Thompson draw
h = (p_target - p_hat) / (p_target if p_hat<=p_target else 1-p_target)
T_x = tau ** h # T_x in [1/tau, tau]
group = policy.rollout(x, n=n, temperature=T_x)
update_counts(x, group, gamma=0.95) # discounted update
PTGS changes only the sampling distribution of training rollouts. The RL update (PPO, Group Relative Policy Optimization (GRPO)) is unchanged, and PTGS is turned off at inference.
The tax is pervasive. Across 42 model-benchmark combinations (14 base/post-trained pairs from Gemma-4, Ministral-3, Qwen2.5, Qwen3.5, times three benchmarks), the calibrated tax at K=128 is positive in 36 of 42 cases. Negative cases cluster at small models on BFCL, where post-training mainly teaches tool-call formatting.
Base models catch up to post-trained ones with enough rollouts. Example from the paper: on WebShop with Gemma-4-31B, base-with-harness reaches over 85% pass@128 versus 56% for the post-trained version. The crossover rollout budget shrinks as models get bigger: Gemma-4 on WebShop crosses at k>128 at 4B but at k≈3 at 31B.
Bimodalization is visible per task. On WebShop with Gemma-4-31B, the “pass given compute” category (sometimes solved) drops from 87.6% to 30.0% after post-training; “always pass” grows from 0% to 26%, but “always fail” also grows from 12.4% to 44%.
The tax is cheaply predictable. The calibrated tax computed from just 8 rollouts on half the tasks correlates at Spearman ρ=0.85 with the tax computed from 32 rollouts on the held-out half. It also predicts the consistency gap and future coverage gaps better than other 8-rollout predictors.
PTGS helps on their own RL runs. Fine-tuning Qwen2.5-7B-Instruct with RAGEN on Sokoban and FrozenLake, PPO+PTGS beats fixed-temperature PPO on both pass@1 and pass@128 while paying a smaller tax. On Sokoban PPO: pass@1 goes 46.5→61.1, pass@128 goes 55.0→69.7, tax goes 0.094→0.081. Ablations show that globally raising temperature or adding top-p/top-k truncation does not reproduce PTGS’s joint gain.
Caveat on causation: the 42-pair analysis is observational. The authors don’t know the exact post-training recipes of open checkpoints, so “RL post-training” is shorthand for whatever mix of SFT, DPO, and RL each lab shipped.
•
If you run many rollouts per task with a verifier, don’t assume the instruct/chat model is the right policy. Test the base model with a lightweight harness on your actual task. The paper shows base often wins at large K, especially for larger backbones. The gain is real only if you have a reliable success check; without a verifier, higher coverage doesn’t convert to usable outputs.
•
Cheap diagnosis before committing. You can estimate Sharpening Tax from ~8 rollouts per task on a modest task sample and extrapolate. If it’s strongly positive on your workload, that’s evidence your post-trained model is giving up reachable solutions.
•
Per-task routing is worth testing. Their routing experiment picks base or post-trained per task based on an 8-rollout tax estimate; it often beats either policy alone at large budgets while keeping pass@1 above the base-only option.
•
If you train your own RL policy with group-based methods (GRPO) or PPO with grouped rollouts, Posterior-Tempered Group Sampling is a drop-in sampler change with no update-rule modification and no extra compute. The authors release code.
•
Turn thinking mode off when comparing coverage per rollout, otherwise post-trained models silently spend more test-time compute per sample and the comparison stops being apples-to-apples. The paper’s thinking-on ablation shows the qualitative conclusion survives either way.
Not evaluated: multimodal models, Vision-Language-Action model systems, open-ended generation where success isn’t binary. Treat those as plausible next experiments, not established.
•
The large-scale analysis is observational: training data and recipes of the open checkpoints are undisclosed, so the paper can’t isolate which post-training step causes sharpening. A controlled replication is listed as future work.
•
The harness advantage is asymmetric. The same scaffolding that helps base models often hurts post-trained models, so “base with harness beats post-trained” compares each checkpoint in its best configuration, not under identical prompting.
•
PTGS gains are demonstrated on two small environments (Sokoban, FrozenLake) with one 7B backbone over 200 training steps. Scaling behavior at frontier-size RL is unverified.
•
Everything hinges on binary task success. For open-ended generation, you’d need quality-gated diversity metrics instead of pass@k.
•
The crossover point depends on model scale and benchmark. For small models on tasks where the base can’t format tool calls at all, post-training genuinely expands coverage and the tax goes negative. Don’t apply the “use the base model” advice to small models or to settings where the base simply can’t produce valid actions.