Classytic

Course

AI Inference Infrastructure: serving a model under real constraints

A generative model in production is a capacity problem before it is anything else. Budget the memory a model and its KV cache actually occupy, batch requests without starving any of them, and scale against demand rather than pretending accelerators are infinite.

38 lessons4h 1mEnglish

What you'll learn

  • Budget AI inference memory and capacity without pretending infrastructure is infinite
  • Separate the memory a model occupies from the memory its KV cache grows into as a conversation lengthens
  • Explain what continuous batching improves, and what it costs the request that was already waiting
  • Separate prompt prefill from autoregressive decode and choose a scheduling policy for the workload
  • Use prefix caching without confusing it with response caching or crossing a tenant trust boundary
  • Choose replicas or tensor parallelism by model fit, independent capacity, communication cost and failure scope
  • Evaluate model routing against explicit quality constraints and optimize cost per useful response
  • Define inference reliability in user terms and choose safe degradation under overload or model failure

Requirements

  • Cloud Infrastructure Foundations, or equivalent understanding of instances, scaling and failure domains
  • No machine learning background. This is about serving a model, not building one

Course content

38 lessons · 4h 1m

About the creator