Israeli researchers eye ChatGPT for genetics
Scientists are developing LLMs trained on biological data to transform how we understand and treat disease.
(TIMES OF ISRAEL) The emergence of artificial intelligence programs like ChatGPT, which can create complex content that resembles human creativity, was made possible by converging huge amounts of data available online with high computational power.
Now Israeli scientists at Sheba Medical Centre have teamed up with US chip giant Nvidia and New York City’s Mount Sinai Hospital to create large language models (LLMs), using the same type of technology that powers applications like ChatGPT, but are trained on the biological language of our bodies, to better understand and treat the diseases that afflict us.
“We want to create a type of ChatGPT of genomics that will allow users to put in a whole genome sequencing of a person and will be capable of answering questions about health risks, or what’s the best drug or treatment for a disease based on the unique genetic makeup of a person,” Avner Halperin, CEO at Sheba Impact at the ARC Centre of Digital Innovation, told The Times of Israel.
“It is a mind-blowing project for truly understanding the force of life that touches at the very core of how the human body works and why each of us is slightly different,” he said.
The three partners are embarking on an ambitious three-year project at an investment of tens of millions of dollars to create a research engine that harnesses generative AI to decode the majority of the human genome – the genetic blueprint for human life – that remains poorly understood.
The brainchild behind the initiative, Professor Gidi Rechavi, head of the Sheba Cancer Research Centre, said that the genomic research engine will be capable of identifying patterns and mechanisms that link a person’s genetic makeup to disease risk and therapeutic response.
The human genome – the set of instructions to build and sustain a human – is made up of 3.2 billion DNA characters or letters. Over the past two decades, science has made advances in the sequencing of a complete human genome, but only two per cent of the human DNA consists of protein-coding genes, while the function of the remainder 98 per cent has been difficult to interpret using traditional approaches, according to Halperin.
The Sheba-led collaboration aims to start decoding the mysteries of the remaining 98 per cent through LLMs and machine learning technologies. Nvidia will provide the computational power and AI infrastructure, and Sheba and Mount Sinai will lend the scientific and clinical expertise to pool and synthesise vast quantities of genomic datasets.
“The human genome is made up of more than 3 billion DNA letters, out of which we understand what two per cent of them do and how they function,” said Halperin. “If a doctor needs to prescribe a drug for depression or administer medicine for a specific cancer, it is mostly via a trial-and-error process, because we haven’t figured out how to read and use 98 per cent of the genomic code.”
The project will bring together clinicians, geneticists, bioinformaticians, and AI researchers, using Nvidia’s computing power, infrastructure and algorithms to link genetic variation to disease risk and therapeutic response across millions of data points.
“AI has the power to unlock the secrets of the human genome and transform health care for billions of people worldwide,” a statement quoted Dr Nati Daniel and Dr Yoli Shavit of Nvidia’s Applied AI Architecture division as saying.
Nvidia’s R&D activities in Israel are the chipmaker’s largest outside of the US. Many of Nvidia’s high-end processors and networking chips, essential for training the largest AI models, are developed at its R&D centres in Israel.
“If successful with this project, we will be able to develop treatments and drugs for the specific genetic needs of people, both for preventive medicine and also for suiting the right intervention,” said Halperin.
“Pharmaceutical companies will use it to develop personalised drugs, hospitals will use it for adapting treatment and choosing drugs with greater accuracy than current methods allow.”