Machine Learning
Use R to investigate how predictive models work. Explore classifiers, trees, ensembles, and clustering before assembling and evaluating a modelling pipeline.
What helps
Basic statistics, tabular data, and algebra are useful. R Foundations provide an entry point if you have not used the language before.
R pathway
Programming Foundations / Practice rooms / Track curriculum and enrollment
Features & the Model Matrix
Build a reusable model-matrix preparation workflow: reserve test rows, fit numeric imputation and scaling, encode categories with saved dictionaries, and apply transformations to new rows in the same feature order. Inspect shape, unknown-category and constant-feature policies. Fit learned preparation inside each training fold so validation and final test information stay outside fitting.
- Train/Test Split: 5 lessons
- Feature Scaling: 5 lessons
- Encoding Categories: 5 lessons
- Transforming Features: 5 lessons
- Missing Data: 5 lessons
k-Nearest Neighbors
Build distances, neighbor ranking, deterministic votes and weighted classification/regression in R. Store training data and fitted preprocessing for reusable prediction, then choose neighbor counts using held-out validation. Explore the effects of scaling, duplicate observations and dimension without assuming one distance rule suits every dataset.
- Distance and Dissimilarity: 5 lessons
- Finding Neighbors: 5 lessons
- The Vote: 5 lessons
- Weighted & Regression kNN: 5 lessons
- Choosing k & Pitfalls: 5 lessons
Naive Bayes
Build Gaussian and multinomial naive Bayes from class priors, fitted feature distributions, fixed vocabularies and add-one smoothing. Use direct log scores to limit underflow and keep class, variance and unknown-token policies explicit. Compose trained continuous-feature and spam-style count classifiers with reusable prediction.
- Bayes' Theorem: 5 lessons
- Gaussian Naive Bayes: 5 lessons
- Multinomial Naive Bayes: 5 lessons
- Smoothing & Log-Space: 5 lessons
- The Full Classifier: 5 lessons
Decision Trees
Build a teaching-scale CART-like numeric classification tree from Gini/entropy, valid threshold search, weighted child impurity and recursive growth. Store feature indices and leaf predictions, support depth and minimum-leaf stops, and perform structural cost-complexity pruning with comparable loss units. Inspect fitted predictions and validate complexity choices on held-out data.
- Impurity: 5 lessons
- Information Gain: 5 lessons
- Growing the Tree: 5 lessons
- Recursive Trees: 5 lessons
- Overfitting & Pruning: 5 lessons
Ensembles
Compose fitted bagged trees and a feature-subsampled forest with bootstrap identities and eligible out-of-bag predictions. Build weighted binary AdaBoost and feature-dependent squared-loss gradient boosting, then explore prediction blending and out-of-fold stacking. Track seeds, stopping policies, signed permutation changes and coverage; ensemble improvement is evaluated rather than guaranteed.
- Bootstrap Aggregation: 5 lessons
- Random Forests: 5 lessons
- AdaBoost: 5 lessons
- Gradient Boosting: 5 lessons
- Combining Models: 5 lessons
k-Means & Partitional Clustering
Build capped Lloyd assignment/update fitting with explicit empty-cluster policies, stored centers and diagnostics, and multiple starts. Distinguish deterministic farthest-first helpers from randomized k-means++ seeding. Evaluate inertia and silhouette and compare compatible fitted results with R kmeans; initialization and local optima limit universal monotonicity claims.
- The Assignment Step: 5 lessons
- The Update Step: 5 lessons
- Seeding & Convergence: 5 lessons
- Choosing k: 5 lessons
- Evaluating Clusters: 5 lessons
Hierarchical & Density Clustering
Build distance/linkage calculations, an agglomerative merge sequence and cuts, and complete DBSCAN core-led expansion with border/noise policies. Compare compatible hierarchical results with R dist/hclust and inspect pairwise partition agreement. These methods offer different geometries, with scale and parameter assumptions rather than guarantees about every cluster shape.
- The Distance Matrix: 5 lessons
- Linkage: 5 lessons
- The Dendrogram: 5 lessons
- DBSCAN: 5 lessons
- Comparing Approaches: 5 lessons
PCA & Dimensionality Reduction
Build PCA from training-fitted centering, optional scaling, sample covariance and eigendecomposition. Store a retained basis for projection and inverse transformation of new rows, inspect reduced reconstruction error and variance retention, and compare aligned results with prcomp. Respect global sign ambiguity, repeated-eigenvalue subspaces, rank deficiency and wide-matrix dimensions.
- Centering & Covariance: 5 lessons
- Eigen-Decomposition: 5 lessons
- Projection & Reconstruction: 5 lessons
- How Many Components: 5 lessons
- Agreement with prcomp: 5 lessons
Model Evaluation & Selection
Build aligned confusion counts, defined metric conventions, complete threshold ROC records and pairwise/trapezoidal AUC checks. Run validation with fold-local fitted preparation and models, preserve candidate identities and select from comparable held-out scores. Refit the chosen candidate before reserved test evaluation and report sample size, class counts and limitations.
- The Confusion Matrix: 5 lessons
- Precision, Recall & F1: 5 lessons
- ROC & AUC: 5 lessons
- Cross-Validation: 5 lessons
- Hyperparameter Tuning: 5 lessons
Capstone: A Predictive Modeling Pipeline
Assemble a complete predictive modeling pipeline with training-fitted imputation and scaling, real kNN/Gaussian naive Bayes/Gini-stump candidates, fold-local fitting and validation-based hyperparameter/model selection. Retain per-candidate held-out predictions and scores, refit the winner on all training rows, and evaluate once on reserved test data with counts, metrics and a training-fitted baseline. Expose fitted state and prediction APIs; retain the one-feature report as a compatible helper rather than the whole system.
- Prepare the Data: 5 lessons
- Train the Contenders: 5 lessons
- Tune by Cross-Validation: 5 lessons
- Final Evaluation: 5 lessons
- The Full Pipeline: 5 lessons