warming up your workspace

Natural Language Processing

Work with language as data in Python. Move from text processing and search to classification, embeddings, and attention, then assemble a small NLP engine.

What helps

An interest in text and language is a good starting point. Later projects use probability, vectors, and comparisons between model predictions.

Python pathway

Programming Foundations / Practice rooms / Track curriculum and enrollment

  1. Text Foundations

    Turn raw text into explicit tokens, a shared vocabulary, IDs and n-grams, then compose per-document corpus statistics. The project teaches normalization as a lossy task choice and preserves document boundaries and fitted feature ordering.

    • Tokenizing: 5 lessons
    • Vocabulary: 5 lessons
    • N-grams: 5 lessons
    • Normalization: 5 lessons
    • Corpus Statistics: 5 lessons
  2. Counting & Vectors

    Once text is tokens, the first model is just counting. This project turns counts into the vectors every classical NLP method uses: word frequencies and the Zipf pattern they follow, bag-of-words vectors over a fixed vocabulary, term-frequency weighting, and similarity between documents. It ends by moving the same operations into numpy, the representation the rest of the track builds on.

    • Word Frequencies: 5 lessons
    • Bag of Words: 5 lessons
    • Term Frequency: 5 lessons
    • Document Similarity: 5 lessons
    • Vectors in numpy: 5 lessons
  3. TF-IDF & Search

    Build document frequency, IDF and TF-IDF, then compare count-cosine retrieval with a fitted smoothed TF-IDF SearchIndex. Queries reuse stored vocabulary and weights; ranking has stable ties and explicit no-match behavior.

    • Inverse Document Frequency: 5 lessons
    • TF-IDF Weighting: 5 lessons
    • Cosine Ranking: 5 lessons
    • A Search Engine: 5 lessons
    • TF-IDF in numpy: 5 lessons
  4. N-gram Language Models

    Estimate bigram probabilities, smooth over a fixed support, score observed events and compute comparable perplexity. Compose a boundary-aware fitted BigramLM with explicit start/end outcomes and bounded sampling. This teaches a count-based language model, not transformer training.

    • Bigram Probabilities: 5 lessons
    • Smoothing: 5 lessons
    • Sentence Probability: 5 lessons
    • Perplexity: 5 lessons
    • Text Generation: 5 lessons
  5. Text Classification

    Sorting text into categories, spam or not, positive or negative, is the workhorse task of applied NLP. This project builds two classic classifiers from scratch: multinomial Naive Bayes, with its class priors and smoothed word likelihoods scored in log space, and logistic regression with the sigmoid and a gradient step. It also builds the metrics that tell you whether a classifier is any good, accuracy, precision, recall, and F1, and wires them into a small sentiment classifier.

    • Naive Bayes Foundations: 5 lessons
    • Classifying with Naive Bayes: 5 lessons
    • Evaluation Metrics: 5 lessons
    • Logistic Regression: 5 lessons
    • A Sentiment Classifier: 5 lessons
  6. Edit Distance & Spelling

    How different are two strings, and what did the user probably mean to type? This project builds edit distance with dynamic programming, the longest common subsequence and a similarity ratio, the four edit operations (insert, delete, replace, transpose) that generate spelling candidates, a frequency-based spell corrector in the style of Norvig's, and fuzzy matching that snaps a misspelling to the nearest known word.

    • Edit Distance: 5 lessons
    • Longest Common Subsequence: 5 lessons
    • Edit Candidates: 5 lessons
    • Spell Correction: 5 lessons
    • Fuzzy Matching: 5 lessons
  7. Sequence Labeling with HMMs

    Many NLP tasks label each word in a sentence: part of speech, named-entity type. The Hidden Markov Model is the classic tool, modeling a hidden chain of tags that each emit a word. This project builds an HMM from counts (transition, emission, and initial distributions), the forward algorithm that sums over all tag sequences to score a sentence, the Viterbi algorithm that finds the single best tag sequence, its application to part-of-speech tagging, and how to evaluate the result, all in numpy.

    • Building the HMM: 5 lessons
    • The Forward Algorithm: 5 lessons
    • The Viterbi Algorithm: 5 lessons
    • Part-of-Speech Tagging: 5 lessons
    • Evaluating a Tagger: 5 lessons
  8. Word Embeddings

    Words that appear in similar contexts have similar meanings, so a word can be represented by the company it keeps. This project builds embeddings from first principles: the co-occurrence matrix counting which words appear near which, the positive pointwise mutual information that turns counts into association strengths, dimensionality reduction with SVD to get dense vectors, cosine similarity between them, and the nearest-neighbor and analogy queries that made embeddings famous.

    • Co-occurrence: 5 lessons
    • Pointwise Mutual Information: 5 lessons
    • Dimensionality Reduction: 5 lessons
    • Vector Similarity: 5 lessons
    • Neighbors & Analogies: 5 lessons
  9. Attention

    Build stable softmax, scaled query/key scores and weighted values with supplied projection matrices. Compose sinusoidal positions, exact optional causal masking, residuals, normalization and feedforward layers into an AttentionEncoder forward pass. Training and text generation are separate outcomes not claimed here.

    • Softmax: 5 lessons
    • Attention Scores: 5 lessons
    • Scaled Dot-Product Attention: 5 lessons
    • Learned Projections: 5 lessons
    • Positional Encoding & the Block: 5 lessons
  10. Capstone: A Mini NLP Engine

    The finale assembles the track into one working system. You will build a preprocessing pipeline that cleans raw text into tokens and a vocabulary, a TF-IDF search that retrieves the most relevant document for a query, a Naive Bayes classifier that labels messages as spam or not, a bigram model that autocompletes the next word, and a dispatcher that routes a request to the right component. By the end you have a small but complete natural-language engine, built from scratch.

    • The Preprocessing Pipeline: 5 lessons
    • Document Search: 5 lessons
    • The Classifier: 5 lessons
    • Autocomplete: 5 lessons
    • The Engine: 5 lessons