Part 5 — Molecules & Biological AI#

⚠️ AI-synthesized; not fully reviewed by course staff. Treat as a study aid; released slides, notebooks, and lecture recordings are authoritative.

Lectures 25-27

What this part is about#

Part 5 is the molecular tier of the course: DNA → proteins/molecules → modern biological AI. The arc is intentional. L25 establishes the population-genetics substrate (variants, GWAS, regulation) that genomic AI builds on. L26 is an honest tour of the classical protein and small-molecule toolkit (sequence alignment, MSAs, PSSMs, fingerprints, docking, QSAR) — the strong baselines that any modern method should be benchmarked against. L27 then tours modern biological AI (ESM-2, AlphaFold, Enformer, Evo, equivariant networks, ProteinMPNN, RFdiffusion) with the running question: do these models actually understand biology, or are they fitting patterns that happen to be biologically interesting?

The answer is “yes, but” — and the qualifier is the load-bearing intellectual content of the part.

Goals for this part#

  • Read a Manhattan plot and a QQ plot; reason about LD and population stratification; explain why polygenic risk scores transfer poorly across ancestries.

  • Build a SMILES → ECFP fingerprint → random-forest baseline; defend a scaffold split; critique a deep-learning paper that claims to “beat classical X.”

  • Critique an ESM-2 / AlphaFold / Evo claim by asking what task and what evaluation. Use pLDDT correctly. Recognize where genomic foundation models still don’t beat task-specific models.

  • Sketch an inverse-folding + experimental-validation pipeline. Treat a generative biology paper as a hypothesis until wet-lab validation arrives.

Key takeaways for this part#

  • Genetic data are high-D, structured, and population-dependent. Population stratification is the single most important confounder.

  • PRS transferability is poor across ancestries — and this matters clinically because it widens, not closes, health disparities.

  • Classical methods (MSA, PSSM, SIFT/PolyPhen, ECFP, docking, QSAR) exploit decades of biological knowledge. They are strong baselines; modern methods compete with them only when they encode similar (or stronger) priors.

  • Modern biological AI re-uses the transformer + NLL + scale recipe in domain costume. Domain constraints (equivariance for 3D molecules, MSAs for protein structure, long context for genomics) still dominate.

  • Generative biology is cheap; experimental validation is the bottleneck. A computational design without wet-lab confirmation is a hypothesis, not a result.

Lectures in this part#

External resources for this part#

Genetics (L25)

  • Visscher et al., “10 Years of GWAS Discovery: Biology, Function, and Translation,” AJHG 101, 2017.

  • Martin et al., “Clinical use of current polygenic risk scores may exacerbate health disparities,” Nat Genet 51, 2019. Required reading.

  • Price et al., “Principal components analysis corrects for stratification in genome-wide association studies,” Nat Genet 38, 2006.

  • GTEx Consortium, “The GTEx Consortium atlas of genetic regulatory effects across human tissues,” Science 369, 2020.

  • ENCODE Project Consortium, “Expanded encyclopaedias of DNA elements,” Nature 583, 2020.

  • Coursera Genomic Data Science (Johns Hopkins) — free, multi-course, hands-on companion.

Proteins & molecules (L26)

  • Berman et al., “The Protein Data Bank,” Nucleic Acids Res 28, 2000.

  • Altschul et al., BLAST, J Mol Biol 215, 1990.

  • Adzhubei et al., PolyPhen-2, Nat Methods 7, 2010 — and Ng & Henikoff, SIFT, Genome Res 11, 2001.

  • Rogers & Hahn, “Extended-Connectivity Fingerprints,” J Chem Inf Model 50, 2010 — Morgan/ECFP fingerprints.

  • Trott & Olson, “AutoDock Vina,” J Comput Chem 31, 2010 — the docking workhorse.

  • Wu et al., “MoleculeNet: a benchmark for molecular machine learning,” Chem Sci 9, 2018.

Modern biological AI (L27)

  • Lin et al., “Evolutionary-scale prediction of atomic-level protein structure” (ESM-2 / ESMFold), Science 379, 2023.

  • Jumper et al., “Highly accurate protein structure prediction with AlphaFold,” Nature 596, 2021.

  • Avsec et al., “Effective gene expression prediction from sequence by integrating long-range interactions” (Enformer), Nat Methods 18, 2021.

  • Nguyen et al., “Sequence modeling and design from molecular to genome scale with Evo,” Science 386, 2024.

  • Dauparas et al., “Robust deep learning–based protein sequence design using ProteinMPNN,” Science 378, 2022.

  • Watson et al., “De novo design of protein structure and function with RFdiffusion,” Nature 620, 2023.

  • Tang & Koo, “Evaluating the representational power of pre-trained DNA language models for regulatory genomics,” Genome Biology 26, 2025 — the genomic-FM critique. Read alongside Enformer / Evo.

AlphaFold DB and resources for hands-on exploration