Reinike AI
Research Paper

Rethinking On-Policy Learning: When Does the Source of AI Data Actually Matter?

Rethinking On-Policy Learning: When Does the Source of AI Data Actually Matter?

In the rapidly evolving world of Large Language Model (LLM) development, a common piece of wisdom has emerged: "on-policy" learning—where a model learns from data it generates itself—is the gold standard. Proponents argue it reduces forgetting, improves generalization, and creates more efficient updates. However, new research from Julianna Piskorz and colleagues at the University of Cambridge and the van der Schaar Lab suggests that the reality is far more nuanced. Their study reveals that the perceived benefits of on-policy data often depend more on how you measure success and which mathematical objectives you choose during training.

Isolating the Rollout Policy

In typical AI training, many factors change at once, making it hard to tell what actually drives performance. The researchers conducted a controlled "strong-to-weak" distillation experiment. They used larger, "teacher" models (like Llama 3 and Qwen 2.5) to train smaller "student" models. Crucially, they isolated the "rollout policy"—the decision of whether the training data comes from the teacher (off-policy) or the student (on-policy). They tested these variables across diverse tasks, including scientific reasoning, medical questions, and complex arithmetic.

The Power of the Objective Function

One of the most significant findings is that the choice of the Kullback–Leibler (KL) divergence—a mathematical measure of how one probability distribution differs from another—matters more than the data source. The study compared "Forward KL" (F-KL) and "Reverse KL" (R-KL). The results showed that Forward KL is remarkably robust; it performs consistently well regardless of whether the data is on-policy or off-policy. In contrast, Reverse KL is much more sensitive and tends to perform significantly better when using on-policy, student-generated data. For business leaders, this means that the "best" training method isn\'t universal—it depends heavily on the specific mathematical objective your engineers are using.

Style Transfer and Incidental Behaviors

The research also explored "style transfer," or how a student model adopts the teacher\'s quirks. In one experiment, they tasked a teacher model to provide reasoning in Spanish. They found that while most setups caused the student to mimic this style, on-policy learning with Reverse KL allowed the student to achieve high task performance without being forced into the teacher’s incidental Spanish style. This suggests that certain training configurations can help models extract "knowledge" without blindly inheriting the "personality" or formatting biases of the source model.

Practical Implications for Enterprise AI

For organizations building or fine-tuning their own AI models, this research offers three practical takeaways. First, don\'t assume that expensive on-policy data collection is always necessary; if using Forward KL, off-policy data from a strong teacher can be just as effective. Second, the learning rate remains the primary driver of "catastrophic forgetting"—the tendency for models to lose old skills when learning new ones. Finally, on-policy data does show a clear advantage in "hard" generalization tasks, such as solving arithmetic problems more complex than those seen during training. The key is to match the training strategy to the specific goal: efficiency, style-matching, or raw reasoning power.