Statistics
Use R to explore data and statistical reasoning. Build from vectors and data frames towards inference, regression, resampling, and a statistical analysis engine.
What helps
Basic arithmetic and graphs help you begin. The later projects use probability and ask you to interpret results alongside their assumptions.
R pathway
Programming Foundations / Practice rooms / Track curriculum and enrollment
R Foundations
Build fluency with R vectors: elementwise arithmetic, recycling, positional and named indexing, missing-value policies, character values and factors. Learn how output type and shape affect later analysis.
- Vectors & Arithmetic: 5 lessons
- Indexing & Subsetting: 5 lessons
- Numbers & Coercion: 5 lessons
- Missing Values: 5 lessons
- Characters & Factors: 5 lessons
Vectors, Lists & the apply Family
Use lists and the apply family to express iteration with deliberate return types. Practice Reduce, Filter, Map, matrices and grouped summaries, while distinguishing concise interfaces from performance guarantees.
- Lists: 5 lessons
- The apply Family: 5 lessons
- Functionals: 5 lessons
- Matrices: 5 lessons
- Split-Apply-Combine: 5 lessons
Data Frames from Scratch
Work with aligned data-frame columns through selection, filtering, derived values, sorting, aggregation and joins. Compose a transaction-value pipeline with explicit row and key behavior.
- Building Data Frames: 5 lessons
- Select & Filter: 5 lessons
- Mutate & Arrange: 5 lessons
- Summarize & Group: 5 lessons
- Joining Frames: 5 lessons
Probability & Distributions
Compute discrete and continuous probabilities, moments and standardized values, then connect formulas with R distribution functions. Seeded simulations illustrate sampling variation, large-sample limits and their assumptions.
- Probability Basics: 5 lessons
- Discrete Distributions: 5 lessons
- Continuous Distributions: 5 lessons
- Expectation & Variance: 5 lessons
- Simulation: 5 lessons
Descriptive Statistics & EDA
Describe center, spread, quantiles, asymmetry and paired relationships. Compare robust summaries and outlier rules while keeping estimator conventions and diagnostic limitations explicit.
- Central Tendency: 5 lessons
- Spread: 5 lessons
- Quantiles & Shape: 5 lessons
- Relationships: 5 lessons
- Outliers & Robustness: 5 lessons
Hypothesis Testing
Build test statistics, p-values and confidence intervals for means and counts, then compare with R test procedures. Examine test assumptions, Type I/II errors, power and multiplicity without equating significance with practical importance.
- Sampling Distributions: 5 lessons
- The t-Test: 5 lessons
- Confidence Intervals: 5 lessons
- Categorical Tests: 5 lessons
- ANOVA & Power: 5 lessons
Linear Regression from Scratch
Derive intercept least-squares fits and matrix formulations, measure residual and in-sample fit quantities, and evaluate predictions. Introduce logistic probabilities and a log-likelihood gradient-ascent update, with identifiability and diagnostic limits.
- Least Squares: 5 lessons
- Measuring Fit: 5 lessons
- The Normal Equations: 5 lessons
- Prediction & Diagnostics: 5 lessons
- Logistic Regression: 5 lessons
Resampling & Simulation
Construct bootstrap and permutation experiments with explicit sampling rules, estimate held-out error through cross-validation, and connect leave-one-out resampling with the jackknife. Distinguish Monte Carlo noise from original-data uncertainty and respect dependence assumptions.
- Resampling Basics: 5 lessons
- Bootstrap Intervals: 5 lessons
- Permutation Tests: 5 lessons
- Cross-Validation: 5 lessons
- Monte Carlo: 5 lessons
Time Series & Smoothing
Study equally spaced time-ordered data through smoothing, lag correlations, differences and a simple additive decomposition. Fit an intercept AR relationship and compare forecast baselines, preserving chronology and avoiding universal stationarity claims.
- Smoothing: 5 lessons
- Autocorrelation: 5 lessons
- Stationarity: 5 lessons
- Decomposition: 5 lessons
- AR Models & Forecasting: 5 lessons
Capstone: A Statistical Analysis Engine
Assemble taught cleaning, exploratory, regression, inference and reporting stages into an inspectable analysis result. Drop missing targets, impute remaining predictor gaps and cap modeling values under an explicit teaching policy. Keep original observed outcomes for the mean test and bootstrap interval; report training fit without claiming held-out validation or causation.
- Cleaning: 5 lessons
- Exploration: 5 lessons
- Modeling: 5 lessons
- Inference: 5 lessons
- The Report: 5 lessons