warming up your workspace

Machine Learning

Use R to investigate how predictive models work. Explore classifiers, trees, ensembles, and clustering before assembling and evaluating a modelling pipeline.

What helps

Basic statistics, tabular data, and algebra are useful. R Foundations provide an entry point if you have not used the language before.

R pathway

Programming Foundations / Practice rooms / Track curriculum and enrollment

  1. Features & the Model Matrix

    Build a reusable model-matrix preparation workflow: reserve test rows, fit numeric imputation and scaling, encode categories with saved dictionaries, and apply transformations to new rows in the same feature order. Inspect shape, unknown-category and constant-feature policies. Fit learned preparation inside each training fold so validation and final test information stay outside fitting.

    • Train/Test Split: 5 lessons
    • Feature Scaling: 5 lessons
    • Encoding Categories: 5 lessons
    • Transforming Features: 5 lessons
    • Missing Data: 5 lessons
  2. k-Nearest Neighbors

    Build distances, neighbor ranking, deterministic votes and weighted classification/regression in R. Store training data and fitted preprocessing for reusable prediction, then choose neighbor counts using held-out validation. Explore the effects of scaling, duplicate observations and dimension without assuming one distance rule suits every dataset.

    • Distance and Dissimilarity: 5 lessons
    • Finding Neighbors: 5 lessons
    • The Vote: 5 lessons
    • Weighted & Regression kNN: 5 lessons
    • Choosing k & Pitfalls: 5 lessons
  3. Naive Bayes

    Build Gaussian and multinomial naive Bayes from class priors, fitted feature distributions, fixed vocabularies and add-one smoothing. Use direct log scores to limit underflow and keep class, variance and unknown-token policies explicit. Compose trained continuous-feature and spam-style count classifiers with reusable prediction.

    • Bayes' Theorem: 5 lessons
    • Gaussian Naive Bayes: 5 lessons
    • Multinomial Naive Bayes: 5 lessons
    • Smoothing & Log-Space: 5 lessons
    • The Full Classifier: 5 lessons
  4. Decision Trees

    Build a teaching-scale CART-like numeric classification tree from Gini/entropy, valid threshold search, weighted child impurity and recursive growth. Store feature indices and leaf predictions, support depth and minimum-leaf stops, and perform structural cost-complexity pruning with comparable loss units. Inspect fitted predictions and validate complexity choices on held-out data.

    • Impurity: 5 lessons
    • Information Gain: 5 lessons
    • Growing the Tree: 5 lessons
    • Recursive Trees: 5 lessons
    • Overfitting & Pruning: 5 lessons
  5. Ensembles

    Compose fitted bagged trees and a feature-subsampled forest with bootstrap identities and eligible out-of-bag predictions. Build weighted binary AdaBoost and feature-dependent squared-loss gradient boosting, then explore prediction blending and out-of-fold stacking. Track seeds, stopping policies, signed permutation changes and coverage; ensemble improvement is evaluated rather than guaranteed.

    • Bootstrap Aggregation: 5 lessons
    • Random Forests: 5 lessons
    • AdaBoost: 5 lessons
    • Gradient Boosting: 5 lessons
    • Combining Models: 5 lessons
  6. k-Means & Partitional Clustering

    Build capped Lloyd assignment/update fitting with explicit empty-cluster policies, stored centers and diagnostics, and multiple starts. Distinguish deterministic farthest-first helpers from randomized k-means++ seeding. Evaluate inertia and silhouette and compare compatible fitted results with R kmeans; initialization and local optima limit universal monotonicity claims.

    • The Assignment Step: 5 lessons
    • The Update Step: 5 lessons
    • Seeding & Convergence: 5 lessons
    • Choosing k: 5 lessons
    • Evaluating Clusters: 5 lessons
  7. Hierarchical & Density Clustering

    Build distance/linkage calculations, an agglomerative merge sequence and cuts, and complete DBSCAN core-led expansion with border/noise policies. Compare compatible hierarchical results with R dist/hclust and inspect pairwise partition agreement. These methods offer different geometries, with scale and parameter assumptions rather than guarantees about every cluster shape.

    • The Distance Matrix: 5 lessons
    • Linkage: 5 lessons
    • The Dendrogram: 5 lessons
    • DBSCAN: 5 lessons
    • Comparing Approaches: 5 lessons
  8. PCA & Dimensionality Reduction

    Build PCA from training-fitted centering, optional scaling, sample covariance and eigendecomposition. Store a retained basis for projection and inverse transformation of new rows, inspect reduced reconstruction error and variance retention, and compare aligned results with prcomp. Respect global sign ambiguity, repeated-eigenvalue subspaces, rank deficiency and wide-matrix dimensions.

    • Centering & Covariance: 5 lessons
    • Eigen-Decomposition: 5 lessons
    • Projection & Reconstruction: 5 lessons
    • How Many Components: 5 lessons
    • Agreement with prcomp: 5 lessons
  9. Model Evaluation & Selection

    Build aligned confusion counts, defined metric conventions, complete threshold ROC records and pairwise/trapezoidal AUC checks. Run validation with fold-local fitted preparation and models, preserve candidate identities and select from comparable held-out scores. Refit the chosen candidate before reserved test evaluation and report sample size, class counts and limitations.

    • The Confusion Matrix: 5 lessons
    • Precision, Recall & F1: 5 lessons
    • ROC & AUC: 5 lessons
    • Cross-Validation: 5 lessons
    • Hyperparameter Tuning: 5 lessons
  10. Capstone: A Predictive Modeling Pipeline

    Assemble a complete predictive modeling pipeline with training-fitted imputation and scaling, real kNN/Gaussian naive Bayes/Gini-stump candidates, fold-local fitting and validation-based hyperparameter/model selection. Retain per-candidate held-out predictions and scores, refit the winner on all training rows, and evaluate once on reserved test data with counts, metrics and a training-fitted baseline. Expose fitted state and prediction APIs; retain the one-feature report as a compatible helper rather than the whole system.

    • Prepare the Data: 5 lessons
    • Train the Contenders: 5 lessons
    • Tune by Cross-Validation: 5 lessons
    • Final Evaluation: 5 lessons
    • The Full Pipeline: 5 lessons