📄 dpo-variants.md

← Vault

DPO Variants

Complete guide to Direct Preference Optimization loss variants in TRL.

Overview

DPO optimizes models using preference data (chosen/rejected pairs). TRL supports 10+ loss variants for different scenarios.

Loss Types

1. Sigmoid (Standard DPO)

Formula: -log(sigmoid(β * logits))

When to use: Default choice, general preference alignment

Config:

`python

DPOConfig(

loss_type="sigmoid",

beta=0.1, # KL penalty

per_device_train_batch_size=64,

learning_rate=1e-6

)

`

2. IPO (Identity Policy Optimization)

Formula: (logits - 1/(2β))²

When to use: Better theoretical foundation, reduce overfitting

Config:

`python

DPOConfig(

loss_type="ipo",

beta=0.1,

per_device_train_batch_size=90,

learning_rate=1e-2

)

`

3. Hinge (SLiC)

Formula: ReLU(1 - β * logits)

When to use: Margin-based objective

Config:

`python

DPOConfig(

loss_type="hinge",

beta=0.1,

per_device_train_batch_size=512,

learning_rate=1e-4

)

`

4. Robust DPO

Formula: Sigmoid with label smoothing for noise robustness

When to use: Noisy preference labels

Config:

`python

DPOConfig(

loss_type="robust",

beta=0.01,

label_smoothing=0.1, # Noise probability

per_device_train_batch_size=16,

learning_rate=1e-3,

max_prompt_length=128,

max_length=512

)

`

5. BCO Pair (Binary Classification)

Formula: Train binary classifier (chosen=1, rejected=0)

When to use: Pairwise preference data

Config:

`python

DPOConfig(

loss_type="bco_pair",

beta=0.01,

per_device_train_batch_size=128,

learning_rate=5e-7,

max_prompt_length=1536,

max_completion_length=512

)

`

6. SPPO Hard

Formula: Push chosen→0.5, rejected→-0.5

When to use: Nash equilibrium, sparse data

Config:

`python

DPOConfig(

loss_type="sppo_hard",

beta=0.1

)

`

7. DiscoPOP

Formula: Log-Ratio Modulated Loss

When to use: Automated loss discovery

Config:

`python

DPOConfig(

loss_type="discopop",

beta=0.05,

discopop_tau=0.05,

per_device_train_batch_size=64,

learning_rate=5e-7

)

`

8. APO Zero

Formula: Increase chosen, decrease rejected likelihood

When to use: Model worse than winning outputs

Config:

`python

DPOConfig(

loss_type="apo_zero",

beta=0.1,

per_device_train_batch_size=64,

learning_rate=2e-7,

max_prompt_length=512,

max_completion_length=512

)

`

9. APO Down

Formula: Decrease both, emphasize rejected reduction

When to use: Model better than winning outputs

Config:

`python

DPOConfig(

loss_type="apo_down",

beta=0.1,

# Same hyperparameters as apo_zero

)

`

10. AOT & AOT Pair

Formula: Distributional alignment via stochastic dominance

When to use: