>

Multimodal & Vision-Language Systems

Combine image and text models: captioning, visual question answering, document understanding and OCR-driven pipelines.

โฑ 6 weeks๐ŸŽš Advanced๐Ÿงฉ 2 projects ยท Capstone๐Ÿ–ฅ Any device ยท our GPUs
Draft syllabus. Timeline and topics are indicative and being finalised with the teaching team โ€” the shape is right, the week-by-week detail may shift.

๐ŸŽฏ Who it's for

Computer-vision or NLP engineers extending into multimodal work.

โœ… Prerequisites

  • Applied Computer Vision or NLP with Transformers
  • Python fluency
  • Comfort with transfer learning

Topics covered

  • Vision-language model families
  • Image-text embeddings and retrieval
  • Captioning and visual question answering
  • Document understanding and OCR pipelines
  • Evaluation for multimodal systems
  • Deployment

Expected completion timeline

6 weeks part-time at 8–12 hours per week ≈ 1.4 months. Self-paced learners can go faster; the live cohort keeps this pace.

Wk 1โ€“2

Vision-language foundations

You produce: An image-text retrieval demo.

Wk 3โ€“4

VQA and document understanding

You produce: A working multimodal pipeline on real documents.

Wk 5โ€“6

Evaluate and capstone

You produce: An evaluated system and a documented deployment.

๐Ÿงญ Where this fits

Advances the Computer Vision / NLP Engineer path toward multimodal roles. See the career & salary map for the full path, the expected pay by region, and the total cost.

๐Ÿ’ณ Delivery & fees

Available in all four delivery formats. Fees vary by format and are confirmed on your advisor call; instalments available. [Add the fee and the next cohort date here.]

Add this to your plan

Book a call and we'll place this course in a full path toward the role you want.

Talk to an advisor