Reinforcement Learning · LLM Reasoning

DelTA vs AntiSD vs DVAO: Three Ways to Fix RLVR's One-Bit Reward

RLVR grades a whole reasoning trace with one bit at the end. DelTA reweights the token update, AntiSD adds a per-token PMI reward, DVAO rebalances several rewards by group variance — three fixes at three layers.

DelTA vs AntiSD vs DVAO: Three Ways to Fix RLVR's One-Bit Reward

The one-bit reward problem

Reinforcement learning from verifiable rewards (RLVR) grades a reasoning trace the way an exam grades an essay: one number at the very end. The answer is right or wrong, and that single scalar is the only feedback the training loop gets. But the parameter update has to act on thousands of individual tokens, so somewhere the algorithm has to decide which tokens actually caused the outcome. Standard GRPO answers this crudely: every token in a correct response is pushed up by the same advantage, every token in a wrong one is pushed down. Commas and lemmas get identical treatment.

That crude answer was good enough to take DeepSeekMath-RL 7B from 46.8% to 51.7% on MATH, and it is still the default in almost every open reasoning recipe. But three 2026 papers attack the same weakness from three different layers of the stack, and choosing between them is a real engineering decision, not a leaderboard preference. DelTA keeps the reward unchanged and reweights the update so credit lands on decisive tokens. AntiSD changes the reward itself, adding a dense per-token signal mined from a privileged context. DVAO leaves the token axis alone entirely and instead fixes how several scalar rewards are combined into one advantage. Same disease, three different organs.

The update layer: DelTA

DelTA’s observation is geometric. A policy-gradient update, it shows, is mathematically equivalent to building a linear discriminator over token-gradient vectors: the update direction is a centroid of advantage-weighted token vectors, and applying it nudges the policy toward the “correct” side of that centroid and away from the “incorrect” side. Naive sequence-level RLVR builds the centroid by plain averaging, which means frequent, low-information tokens dominate it. Formatting tokens, connectives and boilerplate appear in almost every response, correct or not, so they overwhelm the rare tokens that actually separate a right answer from a wrong one.

DelTA estimates per-token coefficients that amplify side-specific directions and down-weight shared ones, then reweights a self-normalized RLVR surrogate with them. The reward is untouched; the same verifiable 0/1 signal drives everything. What changes is which tokens get to spend your gradient budget. The cost is an approximation: to keep coefficient estimation cheap, DelTA uses a layer-restricted token-gradient proxy rather than full-parameter gradients. That is compute you would not otherwise spend, and it is the most likely place the method under-delivers on a different model family.

The payoff, on seven math benchmarks, is 28.40 average accuracy for Qwen3-8B-Base against 25.14 for the strongest of DAPO, SAPO and FIPO (+3.26), and 39.91 against 37.29 on Qwen3-14B-Base (+2.62). Appendix results extend to code generation and an Olmo3-7B-Base run, but the evidence is concentrated on math with Qwen3-Base models.

The reward layer: AntiSD

AntiSD goes further: it builds a new per-token reward. The trick is to invert self-distillation. Ordinary self-distillation with a privileged context (the question plus a hint or partial solution) pulls the policy toward the teacher, and the paper’s diagnostic is that this sharpens exactly the wrong tokens: structural boilerplate gets more confident while deliberation tokens, where the model is weighing options, get flattened. AntiSD instead computes the conditional pointwise mutual information between the next token and the privileged context, and rewards the tokens where the privileged distribution diverges from the unconditioned model. Tokens that carry new information get pushed; tokens that carry none get ignored. Minimizing the gap would be distillation; maximizing it is the “anti.”

Two engineering details keep this from collapsing. The advantage is shaped with a smooth phi function bounded at -0.5·log 2 on the deliberation side, so exploratory tokens cannot be punished too hard. And an entropy gate calibrates a baseline teacher entropy over a 5-step warmup, then deactivates the reward once entropy drops below 0.93 times that baseline, because a confident teacher has no divergence left to mine. The signal is applied only while it is informative, which is why it can be aggressive early and harmless late.

The results are about sample efficiency, not just final accuracy: AntiSD reaches GRPO’s accuracy in 2 to 10 times fewer training steps and finishes up to 11.5 points higher, across five models from 4B to 30B including Qwen3-8B, Qwen3-30B-A3B and two OLMo-3 variants. The strongest single data point is Qwen3-8B on HMMT 2025: GRPO’s peak in roughly one fifth of the steps, ending about 15 points higher. The honest caveat is breadth: a privileged “partial solution” is easy to construct for math and much less obvious for open-ended or agentic tasks.

The objective layer: DVAO

DVAO answers a different question that shows up in real deployments: not “which token deserves credit” but “which reward deserves credit.” Once you want correct and well-formatted and within a length budget and a valid tool call, GRPO-style training needs several scalar rewards merged into one advantage. The two standard mergers both fail in a specific way. Summing rewards first produces large advantage magnitudes that destabilize the gradient. Computing one advantage per objective and adding them with fixed weights is stable but blind: an objective the whole rollout group already satisfies carries almost no learning signal, yet its fixed weight keeps it drowning out the hard objective that needs the gradient.

DVAO reads the empirical reward variance of each objective inside each rollout group and scales that objective’s advantage by it. If all 16 sampled responses already satisfy the format reward, that objective’s variance is near zero and it contributes little; if correctness varies widely across the group, that is where the signal lives. Per-group normalization keeps the combined advantage bounded, and a self-adaptive cross-objective regularizer stops one objective from collapsing another, driven by the same variance statistics rather than a retuned hyperparameter.

On Qwen3-4B-Base with math plus tool-use rewards, DVAO reaches 42.19% average accuracy against 38.99% for reward combination and 38.75% for advantage combination, while pushing length compliance to 99.91% versus roughly 96% for the alternatives. The collapse of the static-weight GDPO baseline to 13.41% in the same setting is the telling line: bad objective weighting does not just underperform, it can fail to train at all. But note the known failure mode of variance weighting itself: an objective that every rollout fails also has low variance, and could be starved exactly when it matters most.

Key numbers

MeasurementDelTAAntiSDDVAOSetting and sourceSame harness?
Where the fix livesToken update reweightingNew per-token rewardObjective-level advantage weightsDesign, all three papersYes, by construction
Extra machinery per stepLayer-restricted gradient proxyPrivileged-context forward passes + entropy gatePer-group variance statisticsEach paperYes
Headline result+3.26 avg (28.40 vs 25.14)+11.5 max, GRPO peak in 2–10x fewer steps42.19% vs 38.99% avg accuracyQwen3-8B-Base math / 4B–30B math / Qwen3-4B-Base math+toolsNo, three different setups
Qwen3-8B math detailAIME24 43.13, AIME25 26.46, HMMT25 18.33HMMT25: GRPO peak in ~1/5 steps, ends ~+15 ptsNot run on this modelDelTA and AntiSD paper tablesNo
Largest model testedQwen3-14B-Base (39.91 vs 37.29, +2.62)Qwen3-30B-A3B (five models 4B–30B)Qwen2.5-7B-InstructEach paperNo
Secondary objectiveNoneEntropy gate at 0.93× warmup baselineLength compliance 99.91% vs ~96%AntiSD / DVAO papersNo
Handles multiple rewardsNo, single verifiable rewardNo, one shaped rewardYes, math + tool use (BFCL-v4)DesignYes
Domains beyond mathCode gains in appendixHumanEval+/MBPP+ as secondary checksBFCL-v4 tool-use suiteEach paperNo

Read the table for what it does not contain: a single row where two of these methods were run under the same data, reward, and evaluation. DelTA, AntiSD and DVAO were each benchmarked against their own chosen baselines on overlapping but distinct suites. The Qwen3-8B math rows come from different papers with different rollout budgets and step counts, so treat the “headline result” row as three separate claims, not a ranking. What is directly comparable is the design row: only DVAO addresses multi-reward training at all, and only AntiSD needs extra forward passes per step.

It is also worth situating these three against the wider GRPO-variant literature. CPPO shows that where you tighten the trust region matters as much as how much: stricter clipping on early tokens plus a prefix-drift budget lifts Qwen3-30B-A3B-Base from 38.19 (GRPO) to 54.79 AIME Avg@16 with the same data and reward. That result does not need a new reward or a new weighting rule at all, which is a reminder that before adding machinery, the clipping policy itself is still a free variable.

When to use which

  • One verifiable reward, and you cannot afford extra forward passes: start with DelTA. It needs nothing at rollout time beyond the standard setup, it keeps the proven response-level reward, and +3.26 average on seven math benchmarks at 8B is a real lift over DAPO-class baselines. The risk is the gradient proxy; budget a quick sanity check on your model family.
  • Rollout generation dominates your training cost and your domain is math-like: AntiSD is the strongest claim of the three, because 2–10x fewer steps is a compute claim, not a leaderboard nudge. The price is a privileged context you must be able to construct, plus an entropy gate with tuned constants (0.93, a 5-step warmup) that deserve a sensitivity check.
  • You train with two or more rewards: DVAO is the only one of the three that applies. Its Qwen3-4B-Base result, top accuracy and 99.91% length compliance in one run, is exactly what fixed-weight schemes struggle to deliver without a per-task sweep. Watch for the starvation failure mode on objectives that fail uniformly.
  • None of the above: if your reward is a learned preference model rather than a verifiable check, all three methods lose their footing. These are RLVR fixes; noisy reward-model RL is a different problem.

Limits and open questions

The obvious gap is the absent controlled comparison: no one has run DelTA, AntiSD and DVAO on identical data, rewards, and compute. Each paper reports its own baselines, its own step counts, and its own suite mixes, so this page deliberately refuses to crown a winner on the numbers. Breadth is a second shared limit: all three headline on math reasoning, where verifiable rewards are cleanest. DelTA’s code and out-of-domain results are appendix-scale; AntiSD’s privileged context is natural for math and undefined for most agentic tasks; DVAO’s tool-use numbers are the thinnest column in its own table. Third, each method carries its own fragility: DelTA’s layer-restricted proxy, AntiSD’s tuned gate constants, DVAO’s uniformly-hard-objective starvation. Finally, none of the three is benchmarked against an exhaustively hand-tuned version of the thing it replaces, be that token weighting, auxiliary losses, or fixed reward coefficients. “Removes the tuning burden” is a promise all three make and none fully cashes.

FAQ

How do DelTA and AntiSD differ if both act per token?

They differ in what they modify. DelTA leaves the reward as the single end-of-answer scalar and reweights the token-level update so decisive tokens dominate the gradient centroid; it needs no new signal, only a gradient proxy. AntiSD builds a new dense reward per token, computed as the pointwise mutual information between the next token and a privileged context, so it needs extra forward passes to evaluate that context but stops relying on outcome credit entirely.

Does AntiSD replace the verifiable reward in RLVR?

No. The PMI reward is an auxiliary per-token signal that shapes exploration; the verifiable outcome reward still decides whether an answer was right. AntiSD’s own ablations would be the place to check the split, but the paper frames the PMI term as a way to reach the same accuracy in 2 to 10 times fewer steps, not as a replacement for verifiable grading.

When does DVAO matter if I train with only one reward?

It does not. DVAO only changes how multiple objectives are merged into one advantage; with a single reward it reduces to ordinary group-relative normalization. Its entire value proposition appears when you combine, say, correctness, format compliance, and a length budget, and fixed weights either destabilize training or drown the hard objective.

Which method should I try first on a Qwen3-8B math run?

If the run is expensive and the domain is math, AntiSD has the strongest efficiency claim on exactly that setup: Qwen3-8B hitting GRPO’s HMMT 2025 peak in about one fifth of the steps and ending roughly 15 points higher. If you want minimal moving parts and no extra forward passes, DelTA’s +3.26 average on seven math benchmarks at the same model size is the conservative pick.