LLMOps & Inference at Scale
Serving frameworks, batching, caching, quantisation, and the cost / latency trade-offs that decide whether an LLM project is viable.
๐ฏ Who it's for
Platform or MLOps engineers running LLM workloads in production.
โ Prerequisites
- MLOps on the Cloud or equivalent
- Comfort with GPUs and containers
- Basic understanding of transformer inference
Topics covered
- Serving frameworks and continuous batching
- KV cache and memory management
- Quantisation and distillation
- Caching and request routing
- Autoscaling GPU fleets
- Cost / latency / quality trade-offs
Expected completion timeline
6 weeks part-time at 8–12 hours per week ≈ 1.4 months. Self-paced learners can go faster; the live cohort keeps this pace.
Serve efficiently
You produce: An LLM endpoint with batching and caching, benchmarked.
Shrink and scale
You produce: A quantised deployment with an autoscaling policy.
Cost model and project
You produce: A cost / latency report and a production-ready config.
๐งญ Where this fits
Advances the MLOps Engineer path toward senior LLMOps roles. See the career & salary map for the full path, the expected pay by region, and the total cost.
๐ณ Delivery & fees
Available in all four delivery formats. Fees vary by format and are confirmed on your advisor call; instalments available.
Add this to your plan
Book a call and we'll place this course in a full path toward the role you want.
Talk to an advisor