Deep Learning
Gradient Checkpointing
Discarding intermediate activations during the forward pass and recomputing them during backward.
Why interviewers ask about it
Trades roughly 30% extra compute for a large activation-memory reduction. The trade works because recompute is compute-bound and storage is memory-bound.
Related terms
This term is part of the free AI/ML Engineer interview preparation module - browse the full glossary for every definition.