Practical NLP: from messy text to a working product

Build a multilingual helpdesk's language features, one working component at a time.

⏰ 10 weeks🎯 Standard🧩 Capstone🖥 Any device · our GPUs
Generated from this course

🎯 About this course

Most NLP courses teach techniques in isolation. This one teaches them as the parts of a single product.

You join Saathi Desk, a fictional helpdesk for small Indian businesses whose customers write in English, Hindi, Punjabi and Hinglish. Each of 32 scenes starts with a problem someone brings you, has you build the solution in code, and shows where it lives in the product, then hands over to the next scene. By the end the pieces fit together into one working, monitored system.

What makes it different

- A running storyline that links every lab, quiz and activity to a product need.
- Honest evaluation from week three: metrics chosen from the cost of mistakes, and a whole scene on data leakage.
- Indian languages as first-class: Unicode, scripts, code-mixing and the tokeniser traps that break Hindi and Punjabi silently.
- Responsible release: bias audits, per-group results, privacy and a release gate.
- Three ways to do every lab: Colab, local Python, or a reading-only path with expected outputs.
- Buy what you need: Foundations, Core, Modern or Full variants.

By the end you will be able to

  • Explain the main NLP task families and choose among them for a product problem.
  • Clean, normalise and tokenise multilingual text, including Devanagari and Gurmukhi, without silently damaging it.
  • Represent documents as vectors and build a search engine whose results can be explained.
  • Build and honestly evaluate a text classifier, choosing the metric from the cost of mistakes and detecting data leakage.
  • Use embeddings for semantic search and topic discovery, and audit them for bias.
  • Implement language models, entity extraction and attention, and explain how recurrent and attention-based models use context.
  • Adopt pretrained transformer models sensibly: transfer learning, safe prompting and retrieval-grounded answers.
  • Release an NLP feature responsibly: per-group evaluation, privacy, hand-off paths, monitoring and an honest presentation.
Curriculum

Topics covered

Working with text 5 topics

  • Scene 0: your first morning, a hands-on warm-up
    Forty-five minutes of clicking and dragging: sort tickets, match meanings, spot a bad test, decide when a bot should hand over. No code, no wrong-answer penalty.
  • What NLP is and where it is used
    What natural language processing is, the handful of task families behind almost every product, and how to tell which one a problem is.
  • Cleaning text and Unicode (including Devanagari and Gurmukhi)
    Why the same-looking text can be different bytes, how Unicode normalisation fixes it, and which 'invisible' characters you must keep.
  • Tokenisation: words, subwords and characters
    How text is cut into pieces a model can count, why subwords won, and what it costs Hindi and Punjabi.
  • Regular expressions and rule-based extraction
    Pattern matching with regular expressions — a fast, explainable way to pull dates, phone numbers and emails out of text.

Representing text 4 topics

  • Bag of words and n-grams
    Turning documents into rows of counts, and why looking at word pairs recovers a little of the order that counting throws away.
  • TF-IDF: weighting words by how informative they are
    Why raw counts over-reward common words, how inverse document frequency fixes it, and how to compute TF-IDF yourself.
  • Similarity and search basics
    Cosine similarity as the standard way to compare text vectors, and how a ranked search result is produced from it.
  • Lab: build a mini search engine
    Put the whole module together: index a small collection and answer queries with ranked, explainable results.

Classical classification 4 topics

  • Naive Bayes: the first text classifier
    How a model that simply multiplies word probabilities can classify spam, and why 'naive' is both its flaw and its strength.
  • Logistic regression and SVM on text
    Linear classifiers on TF-IDF features: the strongest classical baseline, and how to read what they learned.
  • Precision, recall and F1: choosing the right metric
    Why accuracy can lie, how the confusion matrix gives precision, recall and F1, and how to match a metric to the cost of mistakes.
  • Error analysis and data leakage
    How to look at your model's mistakes to improve it, and the leaks that make test scores look better than reality.

Words as vectors 4 topics

  • Meaning from context: how embeddings are built
    The idea that a word is defined by the company it keeps, and how counting co-occurrences and compressing them yields word vectors.
  • Pre-trained embeddings, analogies and bias
    Why most projects start from vectors trained on huge text, how analogies work, and how embeddings quietly absorb society's biases.
  • Document embeddings and semantic search
    Representing whole documents as vectors, finding meaning matches that keyword search misses, and when plain keywords are still the better tool.
  • Topic modelling and clustering
    Finding the themes in a pile of unlabelled documents with NMF and k-means, and judging whether the answer is any good.

Sequences and language models 4 topics

  • N-gram language models and perplexity
    Predicting the next word from the previous few, smoothing for words never seen, and perplexity as the standard score for a language model.
  • Part-of-speech tagging and named entities
    Labelling every word with its role, finding people, places and organisations, and measuring how well you did.
  • RNNs and LSTMs: giving a model a memory
    How a recurrent network reads a sentence one word at a time carrying a hidden state, why it forgets, and how LSTM gates fix that.
  • Sequence-to-sequence models and attention
    Translation as 'read, then write', why squeezing a sentence into one vector fails, and how attention lets the decoder look back at the words that matter.

Transformers 4 topics

  • Attention and the transformer block
    Self-attention lets every word look at every other word at once. See Q, K and V, masking, residual connections and layer normalisation in one block.
  • Pre-trained models and fine-tuning
    Why starting from a model that has read the internet beats starting from scratch, and the options: use as-is, train a small head, or fine-tune.
  • Generative models: prompting and its limits
    How to instruct a generative model reliably, how to handle its messy output in code, and what it cannot be trusted to do.
  • Transformer embeddings and retrieval
    Sentence embeddings from transformers, and the retrieval step behind document-chat assistants: chunk, embed, retrieve, and build a grounded prompt.

Applied NLP 4 topics

  • Information extraction and summarisation
    Pulling structured facts out of text, shortening long documents by choosing their best sentences, and scoring a summary with ROUGE.
  • Question answering and chatbots
    Designing an assistant that answers from trusted content, knows when it does not know, and hands over to a person when it should.
  • Hindi and Punjabi NLP
    What changes when the language is Hindi or Punjabi: scripts and combining marks, tokenisation traps, code-mixing, spelling variation and the resources available.
  • Evaluation, bias and safety
    How to check an NLP system honestly before it reaches people: per-group results, privacy, safety rules, documentation and monitoring.

Capstone and deployment 4 topics

  • Scoping an NLP product
    Turning a vague wish into a small, testable NLP product: the user, the task, the data, the metric and the failure plan.
  • Building the pipeline
    Assemble cleaning, features, model and honest evaluation into one reproducible pipeline.
  • Deploying and monitoring
    Wrapping a model in a safe prediction function, saving it, and watching for drift once real users arrive.
  • Presenting your work
    Telling a decision-maker honestly what your NLP product does, how well, and where it fails, in five minutes.
Course at a glance

See how this course is built

Where the time goes

Longer spokes mean more hours on that module.

Course roadmap click a step

Try a lesson0 / 3

Sample questions from this course. Your progress is kept only in your browser.

Week by week

Drag the slider, or press play, to walk through the course and see what you will have built by each week.

Week 1 of 10
Wk 1

Scenes 1-4: the messy inbox

Task families; script-safe cleaning; tokens and what they cost; the first auto-filled ticket fields.

Wk 2

Scenes 5-8: similarity and the search box

Bag of words, TF-IDF, similar tickets, and a help-centre search engine.

Wk 3

Scenes 9-12: triage you can trust

A baseline, an explainable classifier, the right metric, and finding the leak.

Wk 4

Scenes 13-16: a meaning layer

Embeddings, a bias audit, semantic search and topic discovery.

Wk 5

Scenes 17-18: reply assist and entities

Language models, perplexity and entity extraction.

Wk 6

Scenes 19-22: memory, attention and the transformer

RNN and LSTM intuition, attention, and the transformer block.

Wk 7

Scenes 23-24: adopt, prompt and ground

Transfer learning, safe prompting, retrieval-grounded answers.

Wk 8

Scenes 25-26: summaries and the self-service bot

Extractive summaries, ROUGE, and a bot that knows when to hand over.

Wk 9

Scenes 27-28: Punjabi, Hindi and the release gate

Script-aware tokenisation, code-mixing, per-group evaluation and the launch review.

Wk 10

Scenes 29-32: capstone and presentations

Scoping, pipeline, deployment and monitoring, and the launch presentation.

What you'll build

A helpdesk-style search and triage prototype

The mini search engine (Module 2) and a leak-free triage pipeline (Module 3), ready to show.

Model card and release-gate report

A short document with per-group results, limits, owner and monitoring plan, produced in Modules 7-8.

Capstone: a scoped, tested, deployed NLP feature

A one-page scope, an experiment table, a working pipeline with a guarded predictor, a model card and a five-minute presentation.

Why take this course

  • Every lesson is a scene in one product story, so you always know why you are learning it.
  • Hindi and Punjabi are treated properly, including the tokeniser bugs that fail silently.
  • Evaluation and responsible release are taught from the start, not bolted on at the end.

You will build

  • A script-safe text cleaner and rule-based extractors
  • A help-centre search engine and a triage classifier you can explain
  • Embeddings, a bias audit, semantic search and topic discovery
  • A transformer block, a retrieval pipeline and a guarded predictor with drift alerts
  • A scoped, tested, presented capstone feature

What you can do afterwards

  • The judgement to choose the simplest method that meets the bar, and to say honestly when it does not.
  • A portfolio piece with scope, evidence, model card and presentation.
How to take it

💳 Delivery & fees

Fees are confirmed on an advisor call. Available in all four delivery formats.