Published December 2025 | Version v1
Dissertation Open

Investigating the Epigenetic Code Through Data-Driven Chromosome Structure Modeling

  • 1. University of Chicago

Contributors

Description

The histone code hypothesis proposes that combinatorial patterns of post-translational modifications on histone proteins template three-dimensional genome structure and regulate gene expression. While experimental and computational studies have provided incremental support for this hypothesis, the field has lacked a quantitative framework for rigorously testing the ability of epigenetic sequences to govern 3D genome organization. This work establishes such a framework by developing SCRIBE (Sequence-to-Chromatin structure via epigenetic Representation, Investigation, Benchmarking, and Editing), a coarse-grained polymer model that explicitly represents individual epigenetic marks and employs maximum entropy optimization to parameterize interactions between histone modifications. SCRIBE enables investigation of the histone code hypothesis and the evaluation of potential applications in epigenetic engineering. In Chapter 2, I apply this framework to investigate potential epigenetic therapeutics through \textit{in silico} knock-out and knock-in experiments. Selective perturbations to interaction parameters reveal the range of chromatin conformations accessible through \textit{de novo} epigenetic engineering and identify the individual contributions of each histone modification to genome organization. These simulations demonstrate that certain modifications play disproportionate roles in chromatin architecture, identifying potential targets for therapeutic intervention. In Chapter 3, I establish a quantitative benchmark for evaluating the fundamental limits of chromatin structure prediction from epigenetic sequences. By encoding polymer sequences directly from experimental ChIP-seq data without state-calling algorithms, this approach enables direct assessment of how epigenetic patterns predict chromatin structure. I employ principal component analysis of Hi-C contact matrices to define the theoretical upper bound on model performance, providing an objective ceiling for evaluating any sequence-based prediction approach. Models trained on histone marks successfully capture large-scale compartmentalization, providing quantitative support for the histone code hypothesis. Incorporating transcription factors yields further performance gains, approaching the upper bound and revealing that these additional regulatory mechanisms contribute complementary structural information. Together, these studies establish a foundation for understanding how one-dimensional epigenetic sequences encode three-dimensional genome structure, bridging mechanistic polymer physics with data-driven optimization to enable predictive models for rational design of epigenetic therapeutics.

Files

Thesis (8).pdf

Files (28.6 MB)

Name Size Download all
md5:0283fd68c1179daf0b9178291ee9864b
28.6 MB Preview Download

Additional details

Identifiers

Other
oai:uchicago.tind.io:16570

UChicago Information

Division(s)
Pritzker School of Molecular Engineering