Cracking one of genomics' biggest puzzles: which DNA switches control which genes?
Understanding how genes are controlled is one of the biggest challenges in genomics. Researchers at the University of Copenhagen have now developed a computational method that predicts which regulatory DNA elements control which genes in individual cell types.
Published in Nature Genetics, the work provides researchers with a new tool for interpreting genetic variation and studying gene regulation across human tissues. The method will make it easier to understand how genetic variation influences health and disease and provide a foundation for future biological and medical discoveries.
Although every cell in the human body contains the same DNA, different cell types use different genes. This is controlled in part by regulatory DNA sequences called enhancers, which act as switches that activate specific genes in specific cells.
- Identifying which enhancers regulate which genes, and in which cell types, is one of the major challenges in genomics today. The question matters because most genetic variants associated with common diseases are found in these regulatory regions rather than in genes themselves, explains Robin Andersson, senior author and Associate Professor at the University of Copenhagen.
Learning from single cells
The new machine learning method, called scE2G, uses single-cell genomic data, allowing researchers to predict enhancer-gene interactions separately for each cell type in a sample.
- Until now, mapping enhancer-gene interactions has largely relied on experiments that average signals across millions of pooled cells. Such approaches can overlook important regulatory differences between individual cell types within the same tissue. Since gene regulation often differs between cell types, understanding these differences is essential for studying both development and disease, says Robin Andersson.
To develop the model, the researchers trained it on more than 10,000 experimentally tested enhancer-gene pairs.
One key application of the method is helping researchers interpret genetic variants linked to disease. One example involved a DNA variant associated with lymphocyte counts. The method linked the variant, through long-range regulation, to the genes INPP4B and IL15, providing a plausible explanation for how the variant influences immune cell biology.
The software is freely available and can be applied to new single-cell datasets, making it a resource for researchers studying gene regulation, genetics and human disease.
About the study
The study, Mapping enhancer-gene regulatory interactions from single-cell data, is published in Nature Genetics. It involved researchers from more than a dozen institutions, including the University of Copenhagen, Stanford University, the Broad Institute of MIT and Harvard, Massachusetts General Hospital, the European Molecular Biology Laboratory, and the University of California San Diego. The research was carried out through the Novo Nordisk Foundation Center for Genomic Mechanisms of Disease (NNFC) at the Broad Institute of MIT and Harvard.
Contact
Robin Andersson
Associate Professor
Department of Biology, University of Copenhagen
robin@bio.ku.dk