← Paper page How Sampling Shapes LLM Alignment: From One-Shot Optima to Iterative Dynamics Ask this paper arXiv ↗

How Sampling Shapes LLM Alignment:
From One-Shot Optima to Iterative Dynamics


When a language model learns from people’s comparisons of its answers, does it matter which answers get compared? And what happens when the model keeps learning from its own answers, round after round?

Yurong Chen, Yu He, Michael I. Jordan, Fan Yao · Working paper
Oral presentation, ESIF Economics and AI+ML 2026 · Highlight, FORC 2026

Press → or click to step through

How a model is tuned to human preferences

a prompt the model writes answers candidate answers vs 1 which answers get compared? a person picks the better one trained model reference model 2 stay close to which reference?

Theory usually treats choices 1 and 2 as fixed background. This paper asks how they shape what the model learns.

The lens is IPO (Identity Preference Optimization), a widely used method whose trained model can be written down exactly. The same effects appear for DPO.

Who gets compared decides who comes out on top

who beats whom 60% 60% 90% A B C A beats everyone head-to-head arrow: who usually wins, and how often comparison pool how often each answer is used score: chance of beating a pool answer
33%

With any fixed sampling rule, some preferences make the trained model rank the head-to-head winner below another answer.

Sampling tuned to the preferences, with more weight on better answers, can always put the winner on top.

Illustration with made-up numbers. In IPO with a uniform reference, the trained model gives more probability to answers with higher scores (paper, Prop. 1). Results: Prop. 4 and Thm 5, which also cover a winning group of answers.

Comparing mostly top answers makes the model more peaked

With a strong, consistent ranking, every sampling keeps the right order (Thm 8). How peaked the model gets still depends on sampling.

Even sampling Sampling tilted toward top answers

Tilting comparisons toward the top widens the gaps among the leading answers: sharper choices, less diversity.

Illustration computed with the paper’s formula (3 strong, 5 weak answers, standard rating model). Proved in general when a few strong answers face a long tail of weak ones (Thm 10, Cor. 3.2); DPO behaves alike (App. D). For weaker rankings, even the order can depend on sampling (Prop. 6).

Now the model learns from itself, round after round

current model its answers people compare pairs next model fixed source 1 reference 2 3
  1. 1On-policy shareHow much of the comparison data comes from the current model, rather than from a fixed source.
  2. 2Reference trackingHow closely the reference follows the latest model, rather than a fixed base model.
  3. 3Update strengthHow far each round may move away from the reference.

Each round, the model helps shape its own next training data and its own reference. This mirrors common practice.

Two ways the loop goes wrong, and when it is safe

When preferences go in a circle 1 2 3 like rock–paper–scissors balance point (never reached) The model keeps chasing the cycle and never settles. Worse with stronger updates, more on-policy data, or a reference that follows the latest model.
When preferences form a clear ranking early later much later Probability piles onto the top answer; diversity drains away. A reference that copies the latest model collapses fully. A more diverse reference cannot stop it in the long run.

The safe zone: with a well-anchored reference and updates that are gentle and not too on-policy, training provably settles on one stable answer mix, from any start.

Circling: Prop. 12 (shown for a rock–paper–scissors example). Collapse: Thm 14. Guaranteed settling, for any preferences: Thm 13. Bars are schematic.

Run the loop on the paper’s two example preference sets

who beats whom

Re-computed in your browser from the paper’s update rule and its two example preference sets built from HelpSteer data (App. E.3.1); reproduces Figures 1–2.

The same patterns show up across real human-feedback data

HelpSteer 35,331 answers rated by people Preference sets 4 answers per prompt; each pair judged on one randomly chosen rating (helpfulness, correctness, …) 118 circular sets 4,924 clear-ranking sets oscillation grows as either knob grows diversity drops toward zero as either knob grows

Circling and collapse are not corner cases: they appear broadly in preferences built from real human ratings.

Each set is run through the loop for 3,000 rounds. Paper, Section 5, Figures 3–4 and App. E; simulations of DPO show the same failure modes (App. F).

What this means for alignment pipelines

  1. 1Sampling is a double-edged design choicePreference-aware sampling can restore “better answers rank higher”; tilting it toward top answers makes the model more peaked and less diverse.
  2. 2Learning from your own outputs can backfireWhen the model shapes its own data and reference, training can keep oscillating or collapse onto a single answer.
  3. 3Stability is a design choiceAnchor to a fixed reference, mix in off-policy data, keep updates gentle: the paper gives conditions that guarantee training settles.

Paper page · arXiv · Ask this paper