LLMs
QLoRA
LoRA over a 4-bit NF4-quantised frozen base, with double quantisation and paged optimisers.
Why interviewers ask about it
Saves memory, not compute - typically ~30% slower per step than LoRA. Compute still happens in bf16 after on-the-fly dequantisation.
This term is part of the free AI/ML Engineer interview preparation module - browse the full glossary for every definition.