Deep Learning
Mixed Precision Training
Running matmuls in 16-bit while keeping master weights and optimiser state in fp32.
Why interviewers ask about it
fp16 overflows at 65,504 and needs loss scaling; bf16 has fp32's exponent range and does not. A 7B model needs ~14 GB of weights but ~112 GB to train - optimiser state dominates.
This term is part of the free AI/ML Engineer interview preparation module - browse the full glossary for every definition.