LLMs
RLHF
SFT, then a Bradley–Terry reward model on human preference pairs, then PPO with a KL penalty to the reference policy.
Why interviewers ask about it
It trains on human *rankings*, not human-written answers. The KL term is what prevents reward hacking; omitting it in an explanation is a common tell.
This term is part of the free AI/ML Engineer interview preparation module - browse the full glossary for every definition.