Algorithms for germline genotyping, methylation deconvolution, and somatic phylogeny reconstruction
Files
Publication or External Link
External Link to Data Files
Date
Authors
Advisor
Sahinalp, Suleyman Cenk
Citation
DRUM DOI
Abstract
One of the fundamental problems in nature is to assemble small blocks into known patterns, using a large reference set of known patterns. The complexity of the problem is at least an order of magnitude higher if the reference set is not fully known.
Genotyping germline genes involves assembling sequencing reads into known sequences of genes known as alleles and determining their number of copies in the genomes of individuals. The first part of the thesis presents methods to genotype complex and highly polymorphic germline loci. First, we present ImmunoTyper-SR, an algorithmic approach to genotype and analyze copy numbers of the germline immunoglobulin and T cell receptor genes using Illumina whole genome sequencing (WGS) data. ImmunoTyper-SR is based on a novel combinatorial optimization formulation that aims to minimize the total edit distance between reads and their assigned IGH alleles from a given database, with constraints on the number and distribution of reads across each called allele. We also present Aldy 4, a fast and highly accurate combinatorial optimization method, to genotype a large set of genes responsible for drug metabolism.
The second part of the dissertation focuses on inferring and interpreting data originating from evolutionary processes in somatic cells. DNA methylation is the biological process by which methyl groups are added to certain nucleotides in the DNA that dictate the activity of genes. We present Qombucha, a method for deconvolving bulk DNA methylation data from patient samples into constituent proportions of known cell types. Since the reference cell profiles are only partially known, we also simultaneously fill in the gaps in the methylation values of the representatives of the cell types.
Finally, we present a work in progress of the analysis of phylogenetic evolution of B cell receptor sequences. When the innate immune system fails to destroy pathogenic invaders, various immunoglobulin genes combine to form naive B cell receptors and rapidly accumulate mutations to greatly increase the binding affinity of the antibodies. We survey the existing literature and describe currently available computational tools, their assumptions and limitations and propose various novel combinatorial optimization formulations to solve distinct variants of the B cell receptor phylogeny inference problem.