Although genetic testing can identify the presence of variants associated with diseases such as cancer or diabetes, it often cannot predict the further development of the illness with absolute certainty, especially in the case of common and complex conditions. This gap is becoming one of the key problems in modern genomics.
The United Arab Emirates (UAE) has spent many years creating extensive genetic databases, including the Emirates Genome Programme and a reference genome developed based on 50,000 samples. Now, the focus is shifting to integrating DNA data with other health-influencing information: gene behavior in specific cells, an individual's clinical history, the environment, and lifestyle.
A biobank was opened in Abu Dhabi in April, designed to match biological samples with genomic, clinical, and lifestyle data. Earlier this month, Mohamed bin Zayed University of Artificial Intelligence launched a long-term 'deep phenotyping' study examining the genetic, biological, and behavioral factors underlying the health and diseases of the UAE population.
For Julia Medvedeva, an associate professor of computational biology at MBZUAI, a broader picture is important, as simply reading DNA is only the beginning. Her research focuses on approaches based on single-cell and multi-omics methods to understand gene regulation, origin, and diseases such as type 2 diabetes.
The genome is the complete set of human genetic instructions. Sequencing technologies have made reading these instructions significantly cheaper and faster over the last decade, allowing for the creation of datasets of a scale that was previously impractical. Medvedeva noted that these changes coincided with the emergence of new laboratory techniques and much greater computational power, enabling scientists to study multiple levels of biology simultaneously instead of viewing the genome as the sole answer.
Researchers can now analyze genomic data in parallel with transcriptomics, which studies gene expression, and are increasingly studying individual cells rather than averaging signals across entire tissues. According to Medvedeva, this is critically important because genetics alone does not fully explain the processes occurring in the body. She emphasized that the task is to combine different types of data to understand how genetic differences ultimately affect gene expression, biological pathways, and disease risk.
Single-cell genomics is part of this shift. Instead of analyzing tissue and obtaining an averaged signal from millions of cells, scientists can examine cells individually, identifying those that behave differently. Thus, two people may have similar genetic variants but differ in which genes are active, how their cells respond to the environment, or which biological pathway begins to function incorrectly.
Despite the vast amount of available biological data, Medvedeva cautions against assuming that predictive capabilities have caught up with data collection capabilities. She believes that now is the right time to develop various biological models. Although companies claim good predictions, she notes that they have not yet reached a high level of reliability for some specific predictions.
The field benefits from increased data diversity, large volumes of information, and mathematical methods capable of processing information that was previously inaccessible to researchers. However, translating this capability into reliable predictions for individual patients is still in its early stages.
There is another issue, particularly relevant to the UAE: most of the genetic data used to create disease risk assessment tools historically came from populations of European descent. This matters because a model developed predominantly based on one population may not work equally well when applied elsewhere. For example, polygenic risk scores combine the influence of multiple genetic variants to assess hereditary predisposition to conditions such as diabetes, heart disease, or certain types of cancer, and studies have repeatedly shown that the effectiveness of such scores decreases when applied to groups underrepresented in the original datasets.
Medvedeva discussed this issue during an interview, touching upon attempts to apply genetic screening approaches developed for European populations elsewhere, and the difficulties researchers face when differing baseline populations. She also stressed that the problem is not simply about finding statistical links between a genetic variant and a disease, but about understanding the biological mechanism linking these two phenomena.
This gives UAE genomics programs practical relevance beyond simply building a large database. The Emirates Reference Genome, based on 50,000 samples, identified over 5.2 million previously unknown genetic variants. The broader Emirates Genome Programme aims to collect one million samples across the country.
The point is not that Emirati biology follows different rules. What is important is that medical prediction becomes more reliable when the population being predicted is adequately represented in the evidence base used to create it.
Artificial intelligence enters this field when genomic data becomes too large and complex for traditional analysis. For instance, Google DeepMind's AlphaGenome can analyze DNA sequences up to a million letters long and predict thousands of possible effects on gene regulation. Its Atlas extends this work to billions of possible single-letter changes in human DNA. This is a significant computational step, but it does not mean that AI can reliably predict a person's medical future based on their genome. Medvedeva stated that AI has already changed genomics, but believes that the most important breakthrough may be bringing together people from different disciplines around the right biological questions.
At MBZUAI, she added that collaboration among researchers from different fields helps formulate these questions, develop potential solutions, and then test and validate the created models. Data remains one of the biggest limitations. Genomic and biological datasets can be sparse, noisy, and collected differently across different populations and conditions. Medvedeva concluded that 'we still do not have enough data to fully utilize everything we can do with such data.'
Some areas of genomic medicine are already available to patients in Abu Dhabi. The newborn screening program uses whole-genome sequencing to screen for over 815 treatable childhood genetic diseases, and national programs are also being developed in oncology, pharmacogenomics, rare diseases, and prenatal screening.
However, these applications also demonstrate why genetics needs context. A clear mutation causing a disease can sometimes lead directly to testing or treatment. Predicting a complex condition like type 2 diabetes differs in that genetics interacts with age, environment, behavior, and many other biological factors. This is why Medvedeva is cautious about how genetic information is conveyed to patients. She insists that a predisposition to disease should not be perceived as a guarantee of its development. Instead of sowing fear, genetic risk can point a person toward the need for more regular screening or closer attention to certain aspects of their health. Patients must also understand the meaning of this information, its limitations, and the decisions they can reasonably make based on it. Thus, the promise of genomics is increasingly shifting from telling people what their DNA will say to providing doctors and patients with better information about what might happen, early enough to take action.
