warming up your workspace

Python practice in Natural Language Processing

Browse the rooms before signing in. Opening a room requires an account and follows your existing access. Practice does not issue certificates.

  • Measure the Mix

    Type-token ratio compares distinct spellings with all occurrences. It is useful as a descriptive statistic when samples use comparable lengths and policies.

    stats. Free room.

  • Name Every Word

    A vocabulary defines which token types a vector model can represent. Build it on training text and save it with the model.

    vocab. Free room.

  • Split the Stream

    A consistent cleaner lets training and query text use the same token policy. Here that policy combines splitting, lowercasing and filtering, while accepting their information loss.

    tokenize. Free room.

  • Two at a Time

    Adjacent token pairs capture local order that a vocabulary alone cannot. A bigram model later uses one prior token to predict the next.

    ngrams. Free room.

  • Bag It Up

    A bag-of-words vector gives every document the same feature positions. A model can compare column zero only if that column always means the same token.

    bow. Account access required.

  • How Close?

    Cosine compares directions rather than raw count totals. Positive rescaling leaves it unchanged, but similar token usage still does not guarantee similar meaning.

    similarity. Account access required.

  • Overlap or Not

    Jaccard scales shared types by the size of the combined vocabulary. It can compare sets of different sizes without treating repeated words as additional evidence.

    similarity. Account access required.

  • Rare Is Rich

    A term shared by every document cannot distinguish those documents by presence alone. IDF supplies a corpus-based rarity weight, not a guarantee of relevance.

    idf. Account access required.

  • How Surprised?

    Perplexity transforms average surprise into an effective branching-factor scale. It compares predictive fit only for the same held-out data and tokenization.

    perplexity. Account access required.

  • Smoothing unseen outcomes

    Add-k controls how strongly a count estimate moves toward a uniform distribution. Smaller positive k contributes less artificial mass than add-one.

    smoothing. Account access required.

  • Spam or Not

    A supervised classifier learns from labeled training messages. Word counts must stay separated by class and must not include held-out evaluation messages.

    classify. Account access required.

  • What Comes Next

    A bigram model estimates the next token from one observed context token. The denominator must count opportunities to observe a successor.

    mle. Account access required.

  • Count the Edits

    Levenshtein distance finds the cheapest sequence of insertions, deletions and substitutions. Dynamic programming stores answers for shorter prefixes so each larger answer can reuse them.

    edit. Account access required.

  • Longest Shared

    A longest common subsequence retains character order while allowing gaps. It can reveal shared structure that is not one contiguous substring.

    lcs. Account access required.

  • The Best Guess

    Frequency ranks plausible corrections when context is unavailable. It is a prior preference for common known words, not a guarantee of the writer’s intention.

    corrector. Account access required.

  • Turn Counts to Weight

    Positive PMI retains only above-independence association. Negative PMI means a pair occurs less than independence predicts; it is not simply a synonym for noisy data.

    ppmi. Account access required.

  • Decode the Tags

    A greedy tag at each position can miss the best sequence. Viterbi keeps one best prefix per state and reconstructs its choices backward from the final winner.

    viterbi. Account access required.

  • Look Everywhere

    Scaled dot-product attention combines matching and retrieval. Query/key scores determine weights; changing values alone changes the retrieved content without changing those weights.

    attention. Account access required.

  • Score the Sentence

    Forward inference marginalizes over all possible hidden tag paths. It produces the observation sequence likelihood, not a decoded sequence or posterior tag distribution.

    forward. Account access required.

  • Scores to Weights

    Attention needs nonnegative weights that sum to one. Softmax converts relative scores into that distribution without treating the raw scores as probabilities.

    softmax. Account access required.