Unexpectedly high intrinsic sequence specificity and complex DNA binding of human transcription
factors
Confirmed Presenter: Hamed Najafabadi, McGill University, Canada
Room: Room AD
Moderator(s): Raphaëlle Luisier; Ivo Grosse
Authors List: Show
- Arttu Jolma, University of Toronto, Canada
- Aldo Hernandez-Corchado, McGill University, Canada
- Ally Yang, University of Toronto, Canada
- Ali Fathi, University of Toronto, Canada
- Kaitlin Laverty, University of Toronto, Canada
- Alexander Brechalov, University of Toronto, Canada
- Rozita Razavi, University of Toronto, Canada
- Mihai Albu, University of Toronto, Canada
- Hong Zheng, University of Toronto, Canada
- Ivan Kulakovskiy, Russian Academy of Sciences, Russia
- Hamed Najafabadi, McGill University, Canada
- Tim Hughes, University of Toronto, Canada
Presentation Overview: Show
There is ongoing controversy regarding the degree to which transcription factors (TFs) independently
specify genomic binding: TF binding motifs are short and degenerate, producing excessive binding site
predictions. Here, we present a computational framework, developed in conjunction with large-scale
genomic HT-SELEX (GHT-SELEX), to examine intrinsic sequence-specificities of human TFs.
The first component of this framework is MAGIX, a hierarchical Bayesian model that explicitly captures
the exponential enrichment dynamics of TF-bound genomic DNA fragments across SELEX cycles, producing
quantitative enrichment scores and statistical confidence estimates for genomic binding regions. Applied
to in vitro GHT-SELEX data for 179 TFs across 25 families, we find that genomic binding regions
discovered by MAGIX often display surprisingly high overlap with ChIP-seq peaks for the same TF,
suggesting much higher intrinsic sequence specificity than anticipated.
The second component, RCADEEM, leverages a protein sequence-derived recognition code to model
alternative DNA-binding modes of C2H2 zinc finger (C2H2-ZF) TFs. RCADEEM integrates machine
learning-based prediction of C2H2-ZF sequence preferences with GHT-SELEX/MAGIX binding profiles to infer
which subsets of ZFs are engaged at individual loci. Applied to 86 C2H2-ZF proteins, RCADEEM reveals
that modular, alternative engagement of C2H2-ZF domains is the norm, leading to recognition of multiple
distinct motifs by the same TF. These alternative motifs often evolve by internal duplication and
divergence within the C2H2-ZF array.
Together, this framework reveals that it is common for TFs to delineate a large fraction of their in
vivo genomic binding sites independently of other cellular factors, often through complex,
context-dependent DNA-binding behaviours.
STAN, a computational framework for inferring spatially informed transcription factor activity
Confirmed Presenter: Hatice Ulku Osmanbeyoglu, University of Pittsburgh, United States
Room: Room AD
Moderator(s): Raphaëlle Luisier; Ivo Grosse
Authors List: Show
- Linan Zhang, Ningbo University, China
- April Sagan, University of Pittsburgh, United States
- Bin Qin, University of Pittsburgh, United States
- Haoyu Wang, University of Pittsburgh, United States
- Elena Kim, University of Pittsburgh, United States
- Baoli Hu, University of Pittsburgh, United States
- Hatice Ulku Osmanbeyoglu, University of Pittsburgh, United States
Presentation Overview: Show
Transcription factors (TFs) orchestrate cellular responses to environmental signals and intercellular
communication. The activity of TFs is influenced by neighboring cells, impacting cellular fate and
function. Spatial transcriptomics (ST) allows for the mapping of mRNA expression across tissue samples,
providing insights into the local microenvironment. However, the potential of ST data to systematically
infer TF activity and its role in cell identity has not been fully exploited. We introduce STAN
(Spatially informed Transcription factor Activity Network), a linear mixed-effects computational
approach that predicts spatially informed, spot-specific TF activities by integrating curated
TF–target gene priors, mRNA expression, spatial coordinates, and histological features. We demonstrate
the utility of STAN on lymph node, dorsolateral prefrontal cortex, breast cancer, and glioblastoma ST
datasets, identifying TFs associated with specific cell types, spatial regions, pathological zones, and
ligand–receptor pairs. STAN enhances the utility of ST data, revealing the intricate interplay between
TFs and spatial organization in diverse biological contexts.
From enhancer-scale to locus-scale: modelling gene regulation with sequence-to-function models
Confirmed Presenter: Casper H. Blaauw, VIB.AI, KU Leuven & Hubrecht Institute,
Belgium
Room: Room AD
Moderator(s): Raphaëlle Luisier; Ivo Grosse
Authors List: Show
-
Casper H. Blaauw, VIB.AI, KU Leuven & Hubrecht Institute, Belgium
- Niklas Kempynck, VIB-KU Leuven, Belgium
- Seppe De Winter, VIB-KU Leuven, Belgium
- Vasilieios Konstantakos, VIB-KU Leuven, Belgium
- Eren Can EkÅŸi, VIB-KU Leuven, Belgium
- Sam Dieltiens, VIB-KU Leuven, Belgium
- Darina Abaffyová, VIB-KU Leuven, Belgium
- Valérie Bercier, VIB-KU Leuven, Belgium
- Ibrahim Taskiran, Illumina, United States
- Gert Hulselmans, VIB-KU Leuven, Belgium
- Valerie Christiaens, VIB-KU Leuven, Belgium
- Ludo Van Den Bosch, VIB-KU Leuven, Belgium
- Lukas Mahieu, VIB-KU Leuven, Belgium
- Alexander van Oudenaarden, Hubrecht Institute, Netherlands
- Oliver Hobert, Columbia University, United States
- David R. Kelley, Calico Labs, United States
- Stein Aerts, VIB-KU Leuven, Belgium
Presentation Overview: Show
The main goal of regulatory genomics is to decipher how spatiotemporal regulation of genes is encoded in
the genome. Sequence-to-function models, which learn to link sequence to readouts like chromatin
accessibility or gene expression, are the state-of-the-art for this task. These models primarily operate
at two levels: predicting the scATAC-seq-based chromatin accessibility for individual enhancers and
predicting scRNA-seq-based expression genes based on their surrounding locus.
Here, we present CREsted, a user-friendly package for enhancer modelling covering all steps from
preprocessing to interpretation. Starting from scATAC-seq data, CREsted provides tools to preprocess the
data, train sequence-to-function models, decipher the learned sequence grammar, and design new cell
type-specific enhancers. We showcase CREsted's performance on a variety of tissues and species, ranging
from mouse brain cortex and human cancer states to a whole-organism zebrafish developmental atlas.
Furthermore, we discuss the evolution of the field towards modelling cell type-specific gene
expression-predicting models. Using the toolkit provided in CREsted, we trained a model to predict gene
expression across every cell type of an entire animal, the nematode C. elegans. Using this model, we
extract and cluster high-importance sequence motifs from the genome, which we use to build an
organism-wide motif grammar across tissues. Furthermore, we use in-silico perturbations to link
regulatory elements to genes, and fine-map eQTLs by scoring naturally occurring variants. Finally, we
use atlases from two related nematode species to for cross-species comparison and augmentation.
Altogether, this provides a view into genome regulation at a unique scale.