Algorithms for germline genotyping, methylation deconvolution, and somatic phylogeny reconstruction

dc.contributor.advisorVishkin, Uzien_US
dc.contributor.advisorSahinalp, Suleyman Cenken_US
dc.contributor.authorHari, Ananthen_US
dc.contributor.departmentElectrical Engineeringen_US
dc.contributor.publisherDigital Repository at the University of Marylanden_US
dc.contributor.publisherUniversity of Maryland (College Park, Md.)en_US
dc.date.accessioned2026-07-02T05:34:13Z
dc.date.issued2025en_US
dc.description.abstractOne of the fundamental problems in nature is to assemble small blocks into known patterns, using a large reference set of known patterns. The complexity of the problem is at least an order of magnitude higher if the reference set is not fully known. Genotyping germline genes involves assembling sequencing reads into known sequences of genes known as alleles and determining their number of copies in the genomes of individuals. The first part of the thesis presents methods to genotype complex and highly polymorphic germline loci. First, we present ImmunoTyper-SR, an algorithmic approach to genotype and analyze copy numbers of the germline immunoglobulin and T cell receptor genes using Illumina whole genome sequencing (WGS) data. ImmunoTyper-SR is based on a novel combinatorial optimization formulation that aims to minimize the total edit distance between reads and their assigned IGH alleles from a given database, with constraints on the number and distribution of reads across each called allele. We also present Aldy 4, a fast and highly accurate combinatorial optimization method, to genotype a large set of genes responsible for drug metabolism. The second part of the dissertation focuses on inferring and interpreting data originating from evolutionary processes in somatic cells. DNA methylation is the biological process by which methyl groups are added to certain nucleotides in the DNA that dictate the activity of genes. We present Qombucha, a method for deconvolving bulk DNA methylation data from patient samples into constituent proportions of known cell types. Since the reference cell profiles are only partially known, we also simultaneously fill in the gaps in the methylation values of the representatives of the cell types. Finally, we present a work in progress of the analysis of phylogenetic evolution of B cell receptor sequences. When the innate immune system fails to destroy pathogenic invaders, various immunoglobulin genes combine to form naive B cell receptors and rapidly accumulate mutations to greatly increase the binding affinity of the antibodies. We survey the existing literature and describe currently available computational tools, their assumptions and limitations and propose various novel combinatorial optimization formulations to solve distinct variants of the B cell receptor phylogeny inference problem.en_US
dc.identifierhttps://doi.org/10.13016/nxv2-slea
dc.identifier.urihttp://hdl.handle.net/1903/35823
dc.language.isoenen_US
dc.subject.pqcontrolledBioinformaticsen_US
dc.subject.pqcontrolledComputer scienceen_US
dc.subject.pquncontrolledB cell receptoren_US
dc.subject.pquncontrolledDNA methylationen_US
dc.subject.pquncontrolledGenotypingen_US
dc.subject.pquncontrolledImmunogenomicsen_US
dc.subject.pquncontrolledPharmacogenomicsen_US
dc.titleAlgorithms for germline genotyping, methylation deconvolution, and somatic phylogeny reconstructionen_US
dc.typeDissertationen_US

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Hari_umd_0117E_25808.pdf
Size:
2.79 MB
Format:
Adobe Portable Document Format