>

LLMOps & Inference at Scale

Serving frameworks, batching, caching, quantisation, and the cost / latency trade-offs that decide whether an LLM project is viable.

โฑ 6 weeks๐ŸŽš Advanced๐Ÿงฉ 2 projects ยท Certificate๐Ÿ–ฅ Any device ยท our GPUs
Draft syllabus. Timeline and topics are indicative and being finalised with the teaching team โ€” the shape is right, the week-by-week detail may shift.

๐ŸŽฏ Who it's for

Platform or MLOps engineers running LLM workloads in production.

โœ… Prerequisites

  • MLOps on the Cloud or equivalent
  • Comfort with GPUs and containers
  • Basic understanding of transformer inference

Topics covered

  • Serving frameworks and continuous batching
  • KV cache and memory management
  • Quantisation and distillation
  • Caching and request routing
  • Autoscaling GPU fleets
  • Cost / latency / quality trade-offs

Expected completion timeline

6 weeks part-time at 8–12 hours per week ≈ 1.4 months. Self-paced learners can go faster; the live cohort keeps this pace.

Wk 1โ€“2

Serve efficiently

You produce: An LLM endpoint with batching and caching, benchmarked.

Wk 3โ€“4

Shrink and scale

You produce: A quantised deployment with an autoscaling policy.

Wk 5โ€“6

Cost model and project

You produce: A cost / latency report and a production-ready config.

๐Ÿงญ Where this fits

Advances the MLOps Engineer path toward senior LLMOps roles. See the career & salary map for the full path, the expected pay by region, and the total cost.

๐Ÿ’ณ Delivery & fees

Available in all four delivery formats. Fees vary by format and are confirmed on your advisor call; instalments available. [Add the fee and the next cohort date here.]

Add this to your plan

Book a call and we'll place this course in a full path toward the role you want.

Talk to an advisor