Multimodal & Vision-Language Systems
Combine image and text models: captioning, visual question answering, document understanding and OCR-driven pipelines.
๐ฏ Who it's for
Computer-vision or NLP engineers extending into multimodal work.
โ Prerequisites
- Applied Computer Vision or NLP with Transformers
- Python fluency
- Comfort with transfer learning
Topics covered
- Vision-language model families
- Image-text embeddings and retrieval
- Captioning and visual question answering
- Document understanding and OCR pipelines
- Evaluation for multimodal systems
- Deployment
Expected completion timeline
6 weeks part-time at 8–12 hours per week ≈ 1.4 months. Self-paced learners can go faster; the live cohort keeps this pace.
Vision-language foundations
You produce: An image-text retrieval demo.
VQA and document understanding
You produce: A working multimodal pipeline on real documents.
Evaluate and capstone
You produce: An evaluated system and a documented deployment.
๐งญ Where this fits
Advances the Computer Vision / NLP Engineer path toward multimodal roles. See the career & salary map for the full path, the expected pay by region, and the total cost.
๐ณ Delivery & fees
Available in all four delivery formats. Fees vary by format and are confirmed on your advisor call; instalments available.
Add this to your plan
Book a call and we'll place this course in a full path toward the role you want.
Talk to an advisor