Deep Learning
Vanishing Gradients
Exponential decay of gradient magnitude through depth, caused by repeatedly multiplying by Jacobians with small spectral norm.
Why interviewers ask about it
Sigmoid's derivative caps at 0.25, so 20 layers multiply by at most 0.25²⁰ ≈ 10⁻¹². ReLU mitigates; residual connections solve.
Related terms
This term is part of the free AI/ML Engineer interview preparation module - browse the full glossary for every definition.