Computational Biology
Use R to investigate what biological data can support. Work from experimental design and genomic evidence to single-cell, evolutionary, and multi-omics analyses.
What helps
Basic biology and familiarity with experiments are helpful. Pay particular attention to sample identity, replication, and what an analysis can reasonably conclude.
R pathway
Programming Foundations / Practice rooms / Track curriculum and enrollment
Can This Experiment Answer the Question?
Audit a synthetic treatment study before running statistics. The apparent treatment effect becomes less convincing when sample identity, replication, missing measurements, exclusions, and processing batches are examined. Build an experiment auditor that separates invalid input from repairable imbalance and from a design whose treatment effect is mathematically not identifiable.
- Assay Plus Metadata: 5 lessons
- Groups and Replicates: 5 lessons
- Missingness and Impossible Values: 5 lessons
- Batch and Confounding: 5 lessons
- The Pre-analysis Contract: 5 lessons
Where in the Genome Did the Signal Land?
Investigate a synthetic regulatory assay whose peaks arrive as chromosome coordinates rather than tidy rows. Build trustworthy genomic intervals, distinguish overlap from proximity, preserve ambiguous evidence, measure coverage, respect strand-aware promoters, and produce a candidate-assignment table without pretending that proximity proves regulation.
- Coordinates Are Biological Identity: 5 lessons
- An Overlap Is a Relationship: 5 lessons
- From Individual Peaks to Genomic Shape: 5 lessons
- Promoters, Strand, and Proximity: 5 lessons
- Evidence Without a Causal Claim: 5 lessons
Which Genes Changed, and Can We Trust It?
Investigate a synthetic RNA-seq perturbation study without turning a ranked gene list into a biological verdict. Prove sample identity and design rank, normalize count depth, fit an explicit DESeq2 contrast, distinguish multiple-testing evidence from moderated effect size, and publish a complete gene ledger that preserves filtered and insufficient results.
- The Model Begins at Experimental Design: 5 lessons
- Normalize Before Comparing: 5 lessons
- Fit the Declared Contrast: 5 lessons
- Multiplicity Is Not Effect Size: 5 lessons
- Publish Evidence, Not a Gene Verdict: 5 lessons
Which Biological Programs Are Overrepresented?
Investigate whether a differential gene set contains more members of declared pathways than expected from the genes that could have been selected. Define the measurable universe, compute one-sided hypergeometric evidence and effect ratios, control multiplicity, expose redundant terms and threshold sensitivity, and publish a pathway dossier that suggests follow-up without claiming a pathway was activated.
- The Background Is Part of the Question: 5 lessons
- Test More Than Expected: 5 lessons
- Control and Deconvolve the Pathway Story: 5 lessons
- Challenge the Pathway Story: 5 lessons
- Publish Pathway Evidence Without a Mechanism Claim: 5 lessons
Which Variants Deserve Human Review?
Investigate a synthetic variant call set without turning software annotations into diagnoses. Validate reference and alternate alleles, expand multiallelic records, preserve sample-level depth and genotype evidence, reconcile many transcript consequences, combine population and clinical assertions cautiously, and publish a complete review dossier with conflicts and missing evidence visible.
- One Row Must Mean One Allele: 5 lessons
- Read Support Is Evidence, Not Certainty: 5 lessons
- One Allele Can Have Many Consequences: 5 lessons
- Frequency and Assertions Need Context: 5 lessons
- Build a Human Review Dossier: 5 lessons
How Did the Microbial Community Change?
Investigate a synthetic microbiome census without confusing read proportions with absolute abundance. Align feature counts, taxonomy, and sample context; audit depth and prevalence; construct relative and log-ratio views; measure within-sample and between-sample diversity; separate location from dispersion in group comparisons; and publish a community dossier that keeps compositional limits and design structure visible.
- Assemble the Community Before Summarizing It: 5 lessons
- Counts, Proportions, and Log Ratios Are Different Views: 5 lessons
- Diversity Within a Sample Has Several Meanings: 5 lessons
- Group Separation Can Mean Location or Spread: 5 lessons
- Publish Community Evidence, Not a Microbial Cause: 5 lessons
Which Cell Populations Are Really in the Sample?
Investigate a synthetic single-cell RNA-seq experiment from raw gene-by-cell counts to cautious population evidence. Assemble a SingleCellExperiment, keep cell-level QC decisions and batch context visible, normalize without rewriting counts, select variable genes, construct reproducible low-dimensional geometry, test cluster stability, summarize markers and donor support, and publish a dossier that never equates an algorithmic cluster with a proven cell type.
- A Cell Atlas Begins With Two Stable Axes: 5 lessons
- Low-Quality Libraries Can Become Convincing Populations: 5 lessons
- Normalize Technical Scale, Then Ask Which Genes Vary: 5 lessons
- A Cluster Is an Algorithmic Partition, Not a Cell Type: 5 lessons
- Markers Support an Annotation, They Do Not Prove an Identity: 5 lessons
Which Evolutionary History Does the Sequence Evidence Support?
Build a cautious phylogenetic reconstruction from aligned sequences. Audit alignment identity, calculate bounded evolutionary distances, reconstruct and inspect a rooted hierarchy, challenge clades by resampling sites, and finish with a molecular-evolution dossier that distinguishes a supported topology from a proven history.
- Gate the Alignment: 5 lessons
- Estimate Evolutionary Distance: 5 lessons
- Reconstruct the Tree: 5 lessons
- Challenge the Clades: 5 lessons
- Publish the Molecular Evolution Dossier: 5 lessons
What Can a Protein Sequence Actually Tell Us?
Investigate protein sequence evidence from alphabet and composition through motifs, domain architecture, homology, biophysical clues, and a bounded function-candidate dossier. Every project artifact keeps source coverage, ambiguity, conflicts, and missing experimental support visible.
- Gate the Protein Evidence: 5 lessons
- Find Motifs and Domains: 5 lessons
- Compare Homologous Proteins: 5 lessons
- Read Biophysical Clues: 5 lessons
- Publish the Protein Evidence Dossier: 5 lessons
Can Multiple Omic Layers Support the Same Cohort Story?
Build a leakage-resistant multi-omics study from sample maps and assay contracts through subject-level preprocessing, block-aware integration, held-out validation, distribution-shift auditing, and a model card. The capstone supports only the question and cohort the evidence can defend.
- Harmonize the Cohort: 5 lessons
- Preprocess Without Leakage: 5 lessons
- Integrate Omic Views: 5 lessons
- Validate the Signal: 5 lessons
- Publish the Multi-omics Dossier: 5 lessons