warming up your workspace

Bioinformatics

Explore biological sequences with Python. Start with DNA, then work through alignment, variants, sequence files, and genome search to characterise an unknown sample.

What helps

Knowing the basic relationship between DNA, RNA, and proteins helps. You can use the programming Foundations before working on the sequence projects.

Python pathway

Programming Foundations / Practice rooms / Track curriculum and enrollment

  1. DNA Basics

    Represent canonical DNA as strings: validate the declared alphabet, count bases, distinguish complements and reverse complements, compute composition and find overlapping motif positions. Sequence patterns are observations, not proof of function.

    • Reading DNA: 5 lessons
    • The Complement: 5 lessons
    • GC Content: 5 lessons
    • Finding Motifs: 5 lessons
    • Capstone: Profile a Sequence: 5 lessons
  2. Transcription & Translation

    Model the RNA counterpart of a coding strand, use the standard genetic code, translate declared frames and scan complete ATG-to-stop candidates on both strands. Retain coordinates and distinguish candidate coding regions from verified genes.

    • Transcription: 5 lessons
    • The Genetic Code: 5 lessons
    • Translation: 5 lessons
    • Reading Frames & ORFs: 5 lessons
    • Capstone: Gene to Protein: 5 lessons
  3. Sequence Composition & Motifs

    Characterize a sequence statistically and find the patterns in it: nucleotide and dinucleotide frequencies, CpG sites, melting temperature, k-mer spectra, exact and approximate motif matching, and building a consensus from a set of related sequences.

    • Composition Statistics: 5 lessons
    • K-mers: 5 lessons
    • Motif Matching: 5 lessons
    • Consensus & Profile: 5 lessons
    • Capstone: Characterize a Sequence: 5 lessons
  4. Pairwise Alignment

    Build edit-distance and scored global/local dynamic-programming matrices, recover deterministic traceback paths, and report identity with an explicit denominator. An optimal score depends on its scoring model and does not prove homology.

    • Edit Distance: 5 lessons
    • Scoring & Dot Plots: 5 lessons
    • Needleman-Wunsch (Global): 5 lessons
    • Smith-Waterman (Local): 5 lessons
    • Capstone: Align Two Genes: 5 lessons
  5. Mutations & Variants

    Compare synthetic prealigned references and samples. Report substitutions, individual and joint codon contexts, partial codons and stop changes, and model explicitly described indels. Keep sequence effects separate from functional or clinical claims.

    • Detecting Variants: 5 lessons
    • Coding Effects: 5 lessons
    • Indels & Frameshifts: 5 lessons
    • Applying Variants: 5 lessons
    • Capstone: Variant Report: 5 lessons
  6. FASTA & FASTQ Parsing

    Real sequence data arrives in FASTA and FASTQ files. Parse FASTA records (header + sequence, possibly multi-line), handle many records at once, parse FASTQ reads with their quality strings, decode Phred quality scores, and filter low-quality reads, the daily bread of working with sequencing data.

    • The FASTA Record: 5 lessons
    • Multi-record FASTA: 5 lessons
    • FASTQ Reads: 5 lessons
    • Quality Scores: 5 lessons
    • Capstone: A QC Report: 5 lessons
  7. Phylogenetics

    Measure prealigned sequence differences, handle the finite JC model domain, and build auditable UPGMA trees with cluster-size weighting and branch lengths. A tree is a model-based hypothesis; the algorithm does not validate a molecular clock.

    • Sequence Distances: 5 lessons
    • The Distance Matrix: 5 lessons
    • UPGMA Clustering: 5 lessons
    • Newick Format: 5 lessons
    • Capstone: Build a Tree: 5 lessons
  8. Population Genetics

    Count diploid genotypes and allele copies, keep missing data explicit, compare observed and expected diversity, and report HWE approximation applicability. Compare consistently weighted populations while retaining undefined monomorphic cases.

    • Allele Frequencies: 5 lessons
    • Hardy-Weinberg Equilibrium: 5 lessons
    • Heterozygosity & Diversity: 5 lessons
    • Population Differentiation (Fst): 5 lessons
    • Capstone: Population Report: 5 lessons
  9. Genome Search & Assembly

    Build positional k-mer indexes, verify whole reads after seeding, and assemble bounded forward-only error-free examples. Unsupported joins remain separate contigs, and each resulting contig is checked against the reference.

    • The K-mer Index: 5 lessons
    • Read Mapping: 5 lessons
    • Read Overlaps: 5 lessons
    • Assembly: 5 lessons
    • Capstone: A Mini-Assembler: 5 lessons
  10. Capstone: Characterize an Unknown Sample

    Build a source-bound dossier for a fictional DNA sample: input validation, composition, complete candidate ORFs, restriction recognition and first-strand cuts, and an explicitly prealigned reference comparison. Preserve inputs and limitations without claiming organism identification or biological function.

    • First Look: 5 lessons
    • Candidate Coding Regions: 5 lessons
    • Sequence Features: 5 lessons
    • Compare to a Reference: 5 lessons
    • The Dossier: 5 lessons