LLMs
DPO (Direct Preference Optimization)
Optimising preference pairs directly using the closed-form optimal KL-constrained policy, eliminating the reward model and the RL loop.
Why interviewers ask about it
The partition function Z(x) cancels because it is identical for chosen and rejected. Two models instead of four, one main hyperparameter instead of twelve - but it is offline and cannot exceed its preference data.
Related terms
This term is part of the free AI/ML Engineer interview preparation module - browse the full glossary for every definition.