How Sampling Shapes LLM Alignment:
From One-Shot Optima to Iterative Dynamics
When a language model learns from people’s comparisons of its answers, does it matter which answers get compared? And what happens when the model keeps learning from its own answers, round after round?
Yurong Chen, Yu He, Michael I. Jordan, Fan Yao · Working paper
Oral presentation, ESIF Economics and AI+ML 2026 · Highlight, FORC 2026
Press → or click to step through
How a model is tuned to human preferences
Theory usually treats choices 1 and 2 as fixed background. This paper asks how they shape what the model learns.
The lens is IPO (Identity Preference Optimization), a widely used method whose trained model can be written down exactly. The same effects appear for DPO.
Who gets compared decides who comes out on top
With any fixed sampling rule, some preferences make the trained model rank the head-to-head winner below another answer.
Sampling tuned to the preferences, with more weight on better answers, can always put the winner on top.
Illustration with made-up numbers. In IPO with a uniform reference, the trained model gives more probability to answers with higher scores (paper, Prop. 1). Results: Prop. 4 and Thm 5, which also cover a winning group of answers.
Comparing mostly top answers makes the model more peaked
With a strong, consistent ranking, every sampling keeps the right order (Thm 8). How peaked the model gets still depends on sampling.
Tilting comparisons toward the top widens the gaps among the leading answers: sharper choices, less diversity.
Illustration computed with the paper’s formula (3 strong, 5 weak answers, standard rating model). Proved in general when a few strong answers face a long tail of weak ones (Thm 10, Cor. 3.2); DPO behaves alike (App. D). For weaker rankings, even the order can depend on sampling (Prop. 6).
Now the model learns from itself, round after round
- 1On-policy shareHow much of the comparison data comes from the current model, rather than from a fixed source.
- 2Reference trackingHow closely the reference follows the latest model, rather than a fixed base model.
- 3Update strengthHow far each round may move away from the reference.
Each round, the model helps shape its own next training data and its own reference. This mirrors common practice.
Two ways the loop goes wrong, and when it is safe
The safe zone: with a well-anchored reference and updates that are gentle and not too on-policy, training provably settles on one stable answer mix, from any start.
Circling: Prop. 12 (shown for a rock–paper–scissors example). Collapse: Thm 14. Guaranteed settling, for any preferences: Thm 13. Bars are schematic.
Run the loop on the paper’s two example preference sets
Re-computed in your browser from the paper’s update rule and its two example preference sets built from HelpSteer data (App. E.3.1); reproduces Figures 1–2.
The same patterns show up across real human-feedback data
Circling and collapse are not corner cases: they appear broadly in preferences built from real human ratings.
Each set is run through the loop for 3,000 rounds. Paper, Section 5, Figures 3–4 and App. E; simulations of DPO show the same failure modes (App. F).
What this means for alignment pipelines
- 1Sampling is a double-edged design choicePreference-aware sampling can restore “better answers rank higher”; tilting it toward top answers makes the model more peaked and less diverse.
- 2Learning from your own outputs can backfireWhen the model shapes its own data and reference, training can keep oscillating or collapse onto a single answer.
- 3Stability is a design choiceAnchor to a fixed reference, mix in off-policy data, keep updates gentle: the paper gives conditions that guarantee training settles.