Natural Language Processing
Work with language as data in Python. Move from text processing and search to classification, embeddings, and attention, then assemble a small NLP engine.
What helps
An interest in text and language is a good starting point. Later projects use probability, vectors, and comparisons between model predictions.
Python pathway
Programming Foundations / Practice rooms / Track curriculum and enrollment
Text Foundations
Turn raw text into explicit tokens, a shared vocabulary, IDs and n-grams, then compose per-document corpus statistics. The project teaches normalization as a lossy task choice and preserves document boundaries and fitted feature ordering.
- Tokenizing: 5 lessons
- Vocabulary: 5 lessons
- N-grams: 5 lessons
- Normalization: 5 lessons
- Corpus Statistics: 5 lessons
Counting & Vectors
Once text is tokens, the first model is just counting. This project turns counts into the vectors every classical NLP method uses: word frequencies and the Zipf pattern they follow, bag-of-words vectors over a fixed vocabulary, term-frequency weighting, and similarity between documents. It ends by moving the same operations into numpy, the representation the rest of the track builds on.
- Word Frequencies: 5 lessons
- Bag of Words: 5 lessons
- Term Frequency: 5 lessons
- Document Similarity: 5 lessons
- Vectors in numpy: 5 lessons
TF-IDF & Search
Build document frequency, IDF and TF-IDF, then compare count-cosine retrieval with a fitted smoothed TF-IDF SearchIndex. Queries reuse stored vocabulary and weights; ranking has stable ties and explicit no-match behavior.
- Inverse Document Frequency: 5 lessons
- TF-IDF Weighting: 5 lessons
- Cosine Ranking: 5 lessons
- A Search Engine: 5 lessons
- TF-IDF in numpy: 5 lessons
N-gram Language Models
Estimate bigram probabilities, smooth over a fixed support, score observed events and compute comparable perplexity. Compose a boundary-aware fitted BigramLM with explicit start/end outcomes and bounded sampling. This teaches a count-based language model, not transformer training.
- Bigram Probabilities: 5 lessons
- Smoothing: 5 lessons
- Sentence Probability: 5 lessons
- Perplexity: 5 lessons
- Text Generation: 5 lessons
Text Classification
Sorting text into categories, spam or not, positive or negative, is the workhorse task of applied NLP. This project builds two classic classifiers from scratch: multinomial Naive Bayes, with its class priors and smoothed word likelihoods scored in log space, and logistic regression with the sigmoid and a gradient step. It also builds the metrics that tell you whether a classifier is any good, accuracy, precision, recall, and F1, and wires them into a small sentiment classifier.
- Naive Bayes Foundations: 5 lessons
- Classifying with Naive Bayes: 5 lessons
- Evaluation Metrics: 5 lessons
- Logistic Regression: 5 lessons
- A Sentiment Classifier: 5 lessons
Edit Distance & Spelling
How different are two strings, and what did the user probably mean to type? This project builds edit distance with dynamic programming, the longest common subsequence and a similarity ratio, the four edit operations (insert, delete, replace, transpose) that generate spelling candidates, a frequency-based spell corrector in the style of Norvig's, and fuzzy matching that snaps a misspelling to the nearest known word.
- Edit Distance: 5 lessons
- Longest Common Subsequence: 5 lessons
- Edit Candidates: 5 lessons
- Spell Correction: 5 lessons
- Fuzzy Matching: 5 lessons
Sequence Labeling with HMMs
Many NLP tasks label each word in a sentence: part of speech, named-entity type. The Hidden Markov Model is the classic tool, modeling a hidden chain of tags that each emit a word. This project builds an HMM from counts (transition, emission, and initial distributions), the forward algorithm that sums over all tag sequences to score a sentence, the Viterbi algorithm that finds the single best tag sequence, its application to part-of-speech tagging, and how to evaluate the result, all in numpy.
- Building the HMM: 5 lessons
- The Forward Algorithm: 5 lessons
- The Viterbi Algorithm: 5 lessons
- Part-of-Speech Tagging: 5 lessons
- Evaluating a Tagger: 5 lessons
Word Embeddings
Words that appear in similar contexts have similar meanings, so a word can be represented by the company it keeps. This project builds embeddings from first principles: the co-occurrence matrix counting which words appear near which, the positive pointwise mutual information that turns counts into association strengths, dimensionality reduction with SVD to get dense vectors, cosine similarity between them, and the nearest-neighbor and analogy queries that made embeddings famous.
- Co-occurrence: 5 lessons
- Pointwise Mutual Information: 5 lessons
- Dimensionality Reduction: 5 lessons
- Vector Similarity: 5 lessons
- Neighbors & Analogies: 5 lessons
Attention
Build stable softmax, scaled query/key scores and weighted values with supplied projection matrices. Compose sinusoidal positions, exact optional causal masking, residuals, normalization and feedforward layers into an AttentionEncoder forward pass. Training and text generation are separate outcomes not claimed here.
- Softmax: 5 lessons
- Attention Scores: 5 lessons
- Scaled Dot-Product Attention: 5 lessons
- Learned Projections: 5 lessons
- Positional Encoding & the Block: 5 lessons
Capstone: A Mini NLP Engine
The finale assembles the track into one working system. You will build a preprocessing pipeline that cleans raw text into tokens and a vocabulary, a TF-IDF search that retrieves the most relevant document for a query, a Naive Bayes classifier that labels messages as spam or not, a bigram model that autocompletes the next word, and a dispatcher that routes a request to the right component. By the end you have a small but complete natural-language engine, built from scratch.
- The Preprocessing Pipeline: 5 lessons
- Document Search: 5 lessons
- The Classifier: 5 lessons
- Autocomplete: 5 lessons
- The Engine: 5 lessons