LLMs
Continuous Batching
Evicting finished sequences and admitting new ones every decode step instead of holding a fixed batch.
Why interviewers ask about it
Static batching wastes ~80% of slot-steps because output lengths are long-tailed - one 1,000-token request makes 31 finished slots spin.
Related terms
This term is part of the free AI/ML Engineer interview preparation module - browse the full glossary for every definition.