warming up your workspace

Data Science

Learn Python through the process of answering questions with data. Clean and explore datasets, build models, and evaluate them in an end-to-end machine learning project.

What helps

Basic arithmetic, tables, and graphs are a useful start. Probability and statistics become more important as you move from exploration to inference and modelling.

Python pathway

Programming Foundations / Practice rooms / Track curriculum and enrollment

  1. Data with pandas

    pandas is the foundation of data work in Python. Master the two core structures: the Series (a labeled 1D array) and the DataFrame (a labeled table). Load data, inspect it, select rows and columns, filter, and compute summary statistics, the everyday vocabulary of every data scientist.

    • The Series: 5 lessons
    • The DataFrame: 5 lessons
    • Selecting & Filtering: 5 lessons
    • Computing & Aggregating: 5 lessons
    • Capstone: A First Analysis: 5 lessons
  2. Data Cleaning & Wrangling

    Real data is messy. Detect and handle missing values, find and drop duplicates, fix wrong data types, standardize inconsistent text, and transform columns with map, apply, and binning. Cleaning is where data scientists spend most of their time, and getting it right is what makes every later analysis trustworthy.

    • Missing Data: 5 lessons
    • Duplicates & Types: 5 lessons
    • Cleaning Text: 5 lessons
    • Transforming Data: 5 lessons
    • Capstone: Clean a Dataset: 5 lessons
  3. Exploratory Data Analysis

    Before modeling, you explore. Compute summary statistics (center and spread), aggregate by group, examine distributions with value counts, and measure relationships between variables with correlation. EDA is how you build intuition for a dataset and discover the patterns worth investigating.

    • Summary Statistics: 5 lessons
    • Group Aggregation: 5 lessons
    • Distributions: 5 lessons
    • Relationships: 5 lessons
    • Capstone: Explore a Dataset: 5 lessons
  4. Data Visualization

    A chart often reveals what a table of numbers hides. Build the four workhorse plots with matplotlib: line charts for trends, bar charts for comparing categories, histograms for distributions, and scatter plots for relationships. Label them clearly and choose the right chart for the question.

    • Line Charts: 5 lessons
    • Bar Charts: 5 lessons
    • Histograms: 5 lessons
    • Scatter Plots: 5 lessons
    • Capstone: Choose the Right Chart: 5 lessons
  5. Statistics & Inference

    Calculate descriptive statistics, sampling uncertainty, illustrative confidence intervals and hypothesis tests. State the assumptions behind each method and distinguish observed effects, statistical evidence and practical decisions.

    • Descriptive Statistics: 5 lessons
    • Sampling & Standard Error: 5 lessons
    • Confidence Intervals: 5 lessons
    • Hypothesis Testing: 5 lessons
    • Capstone: An A/B Test: 5 lessons
  6. Linear Regression

    Apply a linear model, fit coefficients with least squares, evaluate squared errors and R-squared, and learn parameters through finite gradient-descent updates. Combine these steps in a house-price teaching example with explicit units and validation limits.

    • The Linear Model: 5 lessons
    • Least Squares: 5 lessons
    • Evaluating the Fit: 5 lessons
    • Gradient Descent: 5 lessons
    • Capstone: Predict House Prices: 5 lessons
  7. Classification

    Predict categories, not numbers. Build two classifiers from scratch: k-nearest neighbors (classify by the majority vote of the closest points) and logistic regression (a sigmoid model trained with gradient descent). Then evaluate them properly with the confusion matrix, accuracy, precision, recall, and F1.

    • k-Nearest Neighbors: 5 lessons
    • Logistic Regression: 5 lessons
    • The Confusion Matrix: 5 lessons
    • Precision, Recall & F1: 5 lessons
    • Capstone: Classify & Evaluate: 5 lessons
  8. Model Evaluation

    Build train/test partitions and cross-validation, compare model flexibility, and learn the statistical meaning of bias and variance. Finish with training-only candidate selection and a calculated final report, while understanding evaluation limits.

    • Train/Test Split: 5 lessons
    • Cross-Validation: 5 lessons
    • Overfitting & Underfitting: 5 lessons
    • The Bias-Variance Tradeoff: 5 lessons
    • Capstone: Evaluate Properly: 5 lessons
  9. Unsupervised Learning

    Find structure in data that has no labels. Build k-means clustering from scratch (group points by similarity), learn to choose the number of clusters, and implement PCA (principal component analysis) to reduce dimensions while keeping the most variance, the two pillars of unsupervised learning.

    • Clustering Basics: 5 lessons
    • The K-Means Algorithm: 5 lessons
    • Choosing the Number of Clusters: 5 lessons
    • Principal Component Analysis: 5 lessons
    • Capstone: Segment Customers: 5 lessons
  10. Capstone: An End-to-End ML Project

    Work through a small synthetic study-hours dataset: prepare paired observations, explore them, fit on the first six cleaned rows, evaluate reserved rows and return a calculated model and prediction report.

    • Load & Clean: 5 lessons
    • Explore: 5 lessons
    • Train the Model: 5 lessons
    • Evaluate: 5 lessons
    • The Complete Pipeline: 5 lessons