Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective

Han Cui1,2,*, Jianhao Yan2,*, Yun Luo2,*, Hongbo Zhang1,2, Zhizhang Fu2, Yue Zhang2,†

1 Zhejiang University   2 Westlake University
* Equal contribution   † Corresponding author

01A student can learn its teacher’s preferences and still collapse.

City: a route to the diamond

In a city with several routes, the teacher points out one path. The student follows it and collects a blue diamond.
The teacher selects a useful path; the student reaches the goal.

Desert: a route that loops

In the desert, three candidate paths diverge. The teacher selects a tangled, self-crossing route. The student follows its twists and loops repeatedly without reaching the shelter in the distance.
The student follows the teacher’s choice, but keeps going in circles instead of reaching shelter.

In on-policy distillation (OPD), the student writes the answers and the teacher scores them. Our experiments reveal how this feedback can make correct reasoning easier to sample—or amplify repetition the teacher rarely produces itself.

02Does OPD Expand Student's Capability Boundary?

With one attempt, the distilled student is much more accurate. Give the initial student more attempts, and it catches up.

Pass at k on AMC23 and AIME24–26: OPD gains are largest at small sampling budgets and shrink as the initial student gets more attempts.
More attempts change the picture. Pass@k measures whether at least one of k sampled answers is correct. AIME uses up to 256 samples; AMC23 uses up to 1,024.

Look closer: coverage & per-problem success

0validated
OPD-only problems

We checked the apparent “new” solutions.

Additional sampling and manual review resolved every candidate in the targeted AIME audit: either the initial student also found a valid solution, or the OPD solution was invalid.

The audited initial student solved 73 problems; the JustRL- and DeepScaleR-distilled students covered 65 and 63. For candidate OPD-only problems, we extended the initial student’s pool from 256 to 1,024 responses.

Audited problem coverage: no validated problem is solved only after OPD.

Success rates increased on 90.6% of already-solvable problems with JustRL and 84.4% with DeepScaleR.

Per-problem success rates largely lie above the equal-performance diagonal after OPD.

03Reward Hacking in OPD: Convergence to a Biased Reward Signal

The loss looks fine. The answers don’t.

3.1 Collapsed OPD Training Still Converges Normally

Both runs converge and concentrate on responses already likely under the initial student.

Both successful and collapsed runs converge normally; their responses concentrate and remain likely under the initial student.

3.2 The Student Faithfully Fits the Reward Signal, Even When the Teacher Is Hacked

Now consider Qwen3-4B teaching Qwen3-1.7B-Base. Optimization converges smoothly, but the student’s responses become overlong and repetitive.

What the teacher writes vs. what the student becomes

AIME24–26 · Final generations
BehaviorTeacherStudent after OPD
Hits the length limit
2.1%
99.4%
Severe repetition
0.0%
38.0%

The student’s accuracy still rises from 0.2% to 5.4%. Accuracy alone misses the degeneration.

Meanwhile, the student increasingly favors the responses preferred by its teacher. The feedback is being learned—even as generation quality deteriorates.

The selection gap measures how much training favors the teacher’s top-ranked over bottom-ranked initial-student responses. It grows modestly with JustRL and sharply in the collapsed Qwen3 setting.

The preference selection gap increases after training, especially in the collapsed Qwen3 setting.

04Diagnosing and Mitigating Reward Hacking

4.1 The Teacher's Reward Signal Can Be Misaligned with Response Quality

The teacher rarely writes this. Why does it reward it?

The warning signs exist before the first OPD update. Some initial-student answers receive favorable teacher feedback despite being excessively long or repetitive.

Teacher’s own response

A solution that ends.

“Number of grand prize outcomes: 1. … Number of prize outcomes (≥ 2 matches): 115. …

P(Grand Prize | Prize) = 1/115.
Thus, m + n = 1 + 115 = 116.”
Correct · 842 tokens
Teacher-preferred student response

An answer that won’t stop.

“Final Answer: 116 …
Answer: 116 …
Final Answer: 116 …
Answer: 116 …”
365×the same answer line,
until the length limit
Correct answer · 8,192 tokens · Truncated

Real excerpts from the paper’s Qwen3-4B examples. Ellipses indicate omitted text. The student example predates OPD; examples illustrate behavior, not prevalence.

Another example: repeated reconsideration
“17,499,840/661,680 = 26.5. Wait, this is incorrect. The calculation is not correct. Instead, we should compute the number of valid sequences. Let’s go back. … But this is not correct. Wait, for the last word to be GY, … Wait, this is not correct.”

Teacher-preferred initial-student response · Incorrect · 7,311 tokens.

What does the teacher favor?

Sort initial-student answers from least to most preferred by the teacher. In the successful setting, preferred answers are more often correct. In the collapsed setting, the highest-ranked answers are disproportionately repetitive.

4.2 Candidate-Pool Interventions Mitigate Reward Hacking

Same teacher. A different trajectory.

Keep the teacher fixed. Change which student responses are available for reinforcement.

A

Mask truncated responses

Set the loss to zero for answers that reach the generation limit.

Recovers from degeneration ↗
B

Start from SFT

Initialize with an SFT student that rarely samples these pathological answers.

Avoids this collapse ↗
Training trajectories of response length, truncation, and repetition: masking settings recover, while the SFT warmup setting avoids the failure mode.
The cycle can be interrupted. Masking recovers as training progresses; SFT initialization avoids the same failure mode from the start.
+3.08percentage points

Average accuracy with masking
10.16% → 13.24%Qwen3-4B teacher · AMC23 & AIME24–26

All teacher sizes & benchmark results
Accuracy (%). Average across four benchmarks.
SettingAMC23AIME24AIME25AIME26Avg.
Base, before OPD2.411.670.001.671.44
4B → Base29.823.333.334.1710.16
+ masking34.648.335.834.1713.24
8B → Base28.315.003.335.0010.41
+ masking32.835.002.503.3310.92
30B-A3B → Base28.616.672.504.1710.49
+ masking33.1310.004.175.0013.07
4B → SFT (warmup)38.869.1710.007.5016.38

Masking improves the average across all three teacher sizes, though not every benchmark improves. Warmup starts from a stronger SFT initialization, so its result includes that initialization advantage.

05Conclusion

Look beyond what the teacher can generate.

Ask how reliably it evaluates what the student actually writes.

OPD amplifies student behaviors through the teacher’s implicit feedback. That perspective connects its gains, its collapse, and the interventions that help.

Scope of these findings

Our experiments focus on small language models and competition mathematics. Coverage is measured with finite sampling and targeted manual audits. Masking and SFT warmup address the overlong, repetitive behavior studied here; these results do not establish a universal solution to OPD collapse.

Successful settings use JustRL-1.5B and DeepScaleR-1.5B-Preview teaching DeepSeek-R1-Distill-Qwen-1.5B on DeepMath. The collapsed setting uses Qwen3-4B teaching Qwen3-1.7B-Base on OpenThoughts3. The paper provides full protocols.

Citation

@misc{cui2026gains,
  title  = {Gains and Collapse in On-Policy Distillation:
            A Reinforcement Learning Perspective},
  author = {Han Cui and Jianhao Yan and Yun Luo and
            Hongbo Zhang and Zhizhang Fu and Yue Zhang},
  year   = {2026},
  url    = {https://github.com/HancCui/opd_hacking}
}