High-Performance Computing
Explore how programs use time and memory with Julia. Work through allocation, parallel patterns, and cache-friendly numerical methods to a high-performance pipeline.
What helps
Start with Julia Foundations if the language is new to you. Familiarity with arrays, numerical calculations, and measuring performance helps with the main projects.
Julia pathway
Programming Foundations / Practice rooms / Track curriculum and enrollment
Type Stability & Performance
Study types, dispatch and inference through typed reductions and generic numerical methods. Distinguish return-type stability from speed, define typed empty identities, and use centered variance rather than subtracting large raw moments. Timing and allocation claims require separate measurements.
- Types & Inference: 5 lessons
- Type-Stable Accumulators: 5 lessons
- Concrete Containers: 5 lessons
- Parametric Methods & Dispatch: 5 lessons
- From Slow to Fast: 5 lessons
Memory & Allocation
Explore dense column-major layout, copy versus view behavior, reusable destinations and broadcast fusion. Observe which storage changes and preserve caller inputs where required. Buffer reuse can reduce allocations without guaranteeing that every operation allocates nothing.
- Array Layout: 5 lessons
- Views vs Copies: 5 lessons
- Preallocation & In-Place: 5 lessons
- Broadcasting & Fusion: 5 lessons
- An Allocation-Light Pipeline: 5 lessons
Reductions & Scans
Build folds, scans, chunk reductions and statistics with explicit identities and order. Include an arbitrary initial value once, preserve noncommutative order where required, and recognize that floating regrouping can change results. Centered updates provide a more robust variance calculation.
- Folds & Reduce: 5 lessons
- Prefix Scans: 5 lessons
- Associativity & Tree Reduction: 5 lessons
- Mapreduce Patterns: 5 lessons
- A Reduction Engine: 5 lessons
SIMD & Vectorization
Write numerical loops suitable for compiler optimization, use valid-index assumptions carefully, and evaluate polynomials with multiply-add expressions. Distinguish muladd from guaranteed fused fma semantics. Correct outputs do not demonstrate generated SIMD instructions or a measured speedup.
- @simd Reductions: 5 lessons
- Branch-Free Code: 5 lessons
- Multiply-Add and Horner Evaluation: 5 lessons
- Vectorized Kernels: 5 lessons
- A Vectorized Pass: 5 lessons
Multithreading
Use real Julia threads for maps and reductions with disjoint chunk-owned state. Balance nonempty chunks, combine generic seeds once, and merge centered statistics. Integer controls can be exact; floating results may vary with grouping. Concurrency safety also depends on callback and aliasing restrictions.
- Parallel Loops: 5 lessons
- Race-Free Reductions: 5 lessons
- Atomics: 5 lessons
- Chunk-Local Accumulation: 5 lessons
- A Parallel Reduction Engine: 5 lessons
Parallel Algorithms
Compose scans, stable compaction, sorted-run merging and parallel chunk sorting. Preserve duplicates, output order and promotion across runs. The examples identify which phases actually use threads; the final map-filter-reduce pipeline retains a sequential filter stage.
- Parallel Scan: 5 lessons
- Parallel Compaction: 5 lessons
- Merging Sorted Runs: 5 lessons
- Parallel Matrix Reductions: 5 lessons
- A Parallel Pipeline: 5 lessons
Tasks & Channels
Coordinate spawned tasks and bounded channels. Preserve result order, wait for all tasks, propagate failures and close streams when downstream work stops. Task completion and deterministic output ordering do not by themselves prevent races or guarantee speed.
- Tasks: 5 lessons
- Channels: 5 lessons
- Producer-Consumer: 5 lessons
- Channel Pipelines: 5 lessons
- A Task-Based Engine: 5 lessons
GPU-Style Kernels
Model one-based grid and block indexing, guarded accesses, grid-stride coverage, gather/scatter and two-level reductions using ordinary Julia arrays on the CPU. Compose a mapped-array and block-reduction engine. Actual GPU execution, memory synchronization and performance require a separate backend and device validation.
- Elementwise Kernels: 5 lessons
- The Grid & Block Model: 5 lessons
- Block Reductions: 5 lessons
- Scatter & Gather: 5 lessons
- A Kernel Engine: 5 lessons
Cache-Friendly Numerics
Implement matrix traversal, tiled multiplication, neighbor stencils and alternative record layouts. Keep boundaries and grid spacing explicit, preserve fractional updates for integer inputs, and distinguish locality reasoning from measured performance. Equivalent floating computations can differ with loop order.
- Matrix Multiply: 5 lessons
- Stencils: 5 lessons
- Structure of Arrays: 5 lessons
- Loop Order & Locality: 5 lessons
- A Numerics Engine: 5 lessons
Capstone: A High-Performance Pipeline
Assemble a numeric-column analytics engine with real threaded chunk execution. A reusable query applies a transform, tests eligibility on the transformed value, accumulates count and sum, then computes a survivor-weighted mean. Preserve empty behavior and type promotion, and distinguish this CPU teaching engine from a benchmarked production analytics system.
- The Dataset: 5 lessons
- The Transform Stage: 5 lessons
- The Filter Stage: 5 lessons
- The Parallel Reduce Stage: 5 lessons
- The Complete Engine: 5 lessons