warming up your workspace

Statistics

Use R to explore data and statistical reasoning. Build from vectors and data frames towards inference, regression, resampling, and a statistical analysis engine.

What helps

Basic arithmetic and graphs help you begin. The later projects use probability and ask you to interpret results alongside their assumptions.

R pathway

Programming Foundations / Practice rooms / Track curriculum and enrollment

  1. R Foundations

    Build fluency with R vectors: elementwise arithmetic, recycling, positional and named indexing, missing-value policies, character values and factors. Learn how output type and shape affect later analysis.

    • Vectors & Arithmetic: 5 lessons
    • Indexing & Subsetting: 5 lessons
    • Numbers & Coercion: 5 lessons
    • Missing Values: 5 lessons
    • Characters & Factors: 5 lessons
  2. Vectors, Lists & the apply Family

    Use lists and the apply family to express iteration with deliberate return types. Practice Reduce, Filter, Map, matrices and grouped summaries, while distinguishing concise interfaces from performance guarantees.

    • Lists: 5 lessons
    • The apply Family: 5 lessons
    • Functionals: 5 lessons
    • Matrices: 5 lessons
    • Split-Apply-Combine: 5 lessons
  3. Data Frames from Scratch

    Work with aligned data-frame columns through selection, filtering, derived values, sorting, aggregation and joins. Compose a transaction-value pipeline with explicit row and key behavior.

    • Building Data Frames: 5 lessons
    • Select & Filter: 5 lessons
    • Mutate & Arrange: 5 lessons
    • Summarize & Group: 5 lessons
    • Joining Frames: 5 lessons
  4. Probability & Distributions

    Compute discrete and continuous probabilities, moments and standardized values, then connect formulas with R distribution functions. Seeded simulations illustrate sampling variation, large-sample limits and their assumptions.

    • Probability Basics: 5 lessons
    • Discrete Distributions: 5 lessons
    • Continuous Distributions: 5 lessons
    • Expectation & Variance: 5 lessons
    • Simulation: 5 lessons
  5. Descriptive Statistics & EDA

    Describe center, spread, quantiles, asymmetry and paired relationships. Compare robust summaries and outlier rules while keeping estimator conventions and diagnostic limitations explicit.

    • Central Tendency: 5 lessons
    • Spread: 5 lessons
    • Quantiles & Shape: 5 lessons
    • Relationships: 5 lessons
    • Outliers & Robustness: 5 lessons
  6. Hypothesis Testing

    Build test statistics, p-values and confidence intervals for means and counts, then compare with R test procedures. Examine test assumptions, Type I/II errors, power and multiplicity without equating significance with practical importance.

    • Sampling Distributions: 5 lessons
    • The t-Test: 5 lessons
    • Confidence Intervals: 5 lessons
    • Categorical Tests: 5 lessons
    • ANOVA & Power: 5 lessons
  7. Linear Regression from Scratch

    Derive intercept least-squares fits and matrix formulations, measure residual and in-sample fit quantities, and evaluate predictions. Introduce logistic probabilities and a log-likelihood gradient-ascent update, with identifiability and diagnostic limits.

    • Least Squares: 5 lessons
    • Measuring Fit: 5 lessons
    • The Normal Equations: 5 lessons
    • Prediction & Diagnostics: 5 lessons
    • Logistic Regression: 5 lessons
  8. Resampling & Simulation

    Construct bootstrap and permutation experiments with explicit sampling rules, estimate held-out error through cross-validation, and connect leave-one-out resampling with the jackknife. Distinguish Monte Carlo noise from original-data uncertainty and respect dependence assumptions.

    • Resampling Basics: 5 lessons
    • Bootstrap Intervals: 5 lessons
    • Permutation Tests: 5 lessons
    • Cross-Validation: 5 lessons
    • Monte Carlo: 5 lessons
  9. Time Series & Smoothing

    Study equally spaced time-ordered data through smoothing, lag correlations, differences and a simple additive decomposition. Fit an intercept AR relationship and compare forecast baselines, preserving chronology and avoiding universal stationarity claims.

    • Smoothing: 5 lessons
    • Autocorrelation: 5 lessons
    • Stationarity: 5 lessons
    • Decomposition: 5 lessons
    • AR Models & Forecasting: 5 lessons
  10. Capstone: A Statistical Analysis Engine

    Assemble taught cleaning, exploratory, regression, inference and reporting stages into an inspectable analysis result. Drop missing targets, impute remaining predictor gaps and cap modeling values under an explicit teaching policy. Keep original observed outcomes for the mean test and bootstrap interval; report training fit without claiming held-out validation or causation.

    • Cleaning: 5 lessons
    • Exploration: 5 lessons
    • Modeling: 5 lessons
    • Inference: 5 lessons
    • The Report: 5 lessons