MLOps
Model Cascade
A cheap model handling confident cases, escalating only the uncertain band to an expensive model.
Why interviewers ask about it
With an 88% early-exit rate, a 2 ms + 180 ms cascade averages ~23 ms. Large latency and cost wins with tunable accuracy loss.
This term is part of the free AI/ML Engineer interview preparation module - browse the full glossary for every definition.