Bioinformatics
Explore biological sequences with Python. Start with DNA, then work through alignment, variants, sequence files, and genome search to characterise an unknown sample.
What helps
Knowing the basic relationship between DNA, RNA, and proteins helps. You can use the programming Foundations before working on the sequence projects.
Python pathway
Programming Foundations / Practice rooms / Track curriculum and enrollment
DNA Basics
Represent canonical DNA as strings: validate the declared alphabet, count bases, distinguish complements and reverse complements, compute composition and find overlapping motif positions. Sequence patterns are observations, not proof of function.
- Reading DNA: 5 lessons
- The Complement: 5 lessons
- GC Content: 5 lessons
- Finding Motifs: 5 lessons
- Capstone: Profile a Sequence: 5 lessons
Transcription & Translation
Model the RNA counterpart of a coding strand, use the standard genetic code, translate declared frames and scan complete ATG-to-stop candidates on both strands. Retain coordinates and distinguish candidate coding regions from verified genes.
- Transcription: 5 lessons
- The Genetic Code: 5 lessons
- Translation: 5 lessons
- Reading Frames & ORFs: 5 lessons
- Capstone: Gene to Protein: 5 lessons
Sequence Composition & Motifs
Characterize a sequence statistically and find the patterns in it: nucleotide and dinucleotide frequencies, CpG sites, melting temperature, k-mer spectra, exact and approximate motif matching, and building a consensus from a set of related sequences.
- Composition Statistics: 5 lessons
- K-mers: 5 lessons
- Motif Matching: 5 lessons
- Consensus & Profile: 5 lessons
- Capstone: Characterize a Sequence: 5 lessons
Pairwise Alignment
Build edit-distance and scored global/local dynamic-programming matrices, recover deterministic traceback paths, and report identity with an explicit denominator. An optimal score depends on its scoring model and does not prove homology.
- Edit Distance: 5 lessons
- Scoring & Dot Plots: 5 lessons
- Needleman-Wunsch (Global): 5 lessons
- Smith-Waterman (Local): 5 lessons
- Capstone: Align Two Genes: 5 lessons
Mutations & Variants
Compare synthetic prealigned references and samples. Report substitutions, individual and joint codon contexts, partial codons and stop changes, and model explicitly described indels. Keep sequence effects separate from functional or clinical claims.
- Detecting Variants: 5 lessons
- Coding Effects: 5 lessons
- Indels & Frameshifts: 5 lessons
- Applying Variants: 5 lessons
- Capstone: Variant Report: 5 lessons
FASTA & FASTQ Parsing
Real sequence data arrives in FASTA and FASTQ files. Parse FASTA records (header + sequence, possibly multi-line), handle many records at once, parse FASTQ reads with their quality strings, decode Phred quality scores, and filter low-quality reads, the daily bread of working with sequencing data.
- The FASTA Record: 5 lessons
- Multi-record FASTA: 5 lessons
- FASTQ Reads: 5 lessons
- Quality Scores: 5 lessons
- Capstone: A QC Report: 5 lessons
Phylogenetics
Measure prealigned sequence differences, handle the finite JC model domain, and build auditable UPGMA trees with cluster-size weighting and branch lengths. A tree is a model-based hypothesis; the algorithm does not validate a molecular clock.
- Sequence Distances: 5 lessons
- The Distance Matrix: 5 lessons
- UPGMA Clustering: 5 lessons
- Newick Format: 5 lessons
- Capstone: Build a Tree: 5 lessons
Population Genetics
Count diploid genotypes and allele copies, keep missing data explicit, compare observed and expected diversity, and report HWE approximation applicability. Compare consistently weighted populations while retaining undefined monomorphic cases.
- Allele Frequencies: 5 lessons
- Hardy-Weinberg Equilibrium: 5 lessons
- Heterozygosity & Diversity: 5 lessons
- Population Differentiation (Fst): 5 lessons
- Capstone: Population Report: 5 lessons
Genome Search & Assembly
Build positional k-mer indexes, verify whole reads after seeding, and assemble bounded forward-only error-free examples. Unsupported joins remain separate contigs, and each resulting contig is checked against the reference.
- The K-mer Index: 5 lessons
- Read Mapping: 5 lessons
- Read Overlaps: 5 lessons
- Assembly: 5 lessons
- Capstone: A Mini-Assembler: 5 lessons
Capstone: Characterize an Unknown Sample
Build a source-bound dossier for a fictional DNA sample: input validation, composition, complete candidate ORFs, restriction recognition and first-strand cuts, and an explicitly prealigned reference comparison. Preserve inputs and limitations without claiming organism identification or biological function.
- First Look: 5 lessons
- Candidate Coding Regions: 5 lessons
- Sequence Features: 5 lessons
- Compare to a Reference: 5 lessons
- The Dossier: 5 lessons