November 5, 2012

Computers 'taught' to ID regulating gene sequences

by Johns Hopkins University School of Medicine

Johns Hopkins researchers have succeeded in teaching computers how to identify commonalities in DNA sequences known to regulate gene activity, and to then use those commonalities to predict other regulatory regions throughout the genome. The tool is expected to help scientists better understand disease risk and cell development.

The work was reported in two recent papers in Genome Research, published online on July 3 and Sept. 27.

"Our goal is to understand how regulatory information is encrypted and to learn which sequence variations contribute to medical risks," says Andrew McCallion, Ph.D., associate professor of molecular and comparative pathobiology in the McKusick-Nathans Institute of Genetic Medicine at Hopkins. "We give data to a computer and 'teach it' to distinguish between data that has no biological value versus data that has this or that biological value. It then establishes a set of rules, which allows it to look at new sets of data and apply what it learned. We're basically sending our computers to school."

These state-of-the-art "machine learning" techniques were developed by Michael Beer, Ph.D., assistant professor of biomedical engineering at the Johns Hopkins School of Medicine, and by Ivan Ovcharenko, Ph.D., at the National Center for Biotechnology Information. The researchers began both studies by creating "training sets" for their computers to "learn" from. These training sets were lists of DNA sequences taken from regions of the genome, called enhancers, that are known to increase the activity of particular genes in particular cells.

For the first of their studies, McCallion's team created a training set of enhancer sequences specific to a particular region of the brain by compiling a list of 211 published sequences that had been shown, by various studies in mice and zebrafish, to be active in the development or function of that part of the brain.

For a second study, the team generated a training set through experiments of their own. They began with a purified population of mouse melanocytes, which are the skin cells that produce the pigment melanin that gives color to skin and absorbs harmful UV rays from the sun. The researchers used a technique called ChIP-seq (pronounced "chip seek") to collect and sequence all of the pieces of DNA that were bound in those cells by special enhancer-binding proteins, generating a list of about 2,500 presumed melanocyte enhancer sequences.

Once the researchers had these two training sets for their computers, one specific to the brain and another to melanocytes, the computers were able to distinguish the features of the training sequences from the features of all other sequences in the genome, and create rules that defined one set from the other. Applying those rules to the whole genome, the computers were able to discover thousands of probable brain or melanocyte enhancer sequences that fit the features of the training sets.

In the brain study, the computers identified 40,000 probable brain enhancer sequences; for melanocytes, 7,500. Randomly testing a subset of each batch of sequences, the scientists found that more than 85 percent of the predicted enhancer sequences enhanced gene activity in the brain or in melanocytes, as expected, verifying the predictive power of their approach.

The researchers say that, in addition to identifying specific DNA sequences that control the genetic activity of a particular organ or cell type, these studies contribute to our understanding of enhancers in general and have validated an experimental approach that can be applied to many other biological questions as well.

More information:
Brain: www.genome.org/cgi/doi/10.1101/gr.139717.112
Melanocyte: www.genome.org/cgi/doi/10.1101/gr.139360.112

Journal information: Genome Research

Provided by Johns Hopkins University School of Medicine

Citation: Computers 'taught' to ID regulating gene sequences (2012, November 5) retrieved 26 April 2024 from https://phys.org/news/2012-11-taught-id-gene-sequences.html

This document is subject to copyright. Apart from any fair dealing for the purpose of private study or research, no part may be reproduced without the written permission. The content is provided for information purposes only.

Explore further

Scientists Map Genetic Regulatory Elements for the Heart

0 shares

Feedback to editors

Computers 'taught' to ID regulating gene sequences

High-precision blood glucose level prediction achieved by few-molecule reservoir computing

Enhancing memory technology: Multiferroic nanodots for low-power magnetic storage

Researchers advance detection of gravitational waves to study collisions of neutron stars and black holes

Automated machine learning robot unlocks new potential for genetics research

AI deciphers new gene regulatory code in plants and makes accurate predictions for newly sequenced genomes

Unveiling a new quantum frontier: Frequency-domain entanglement

Study details a common bacterial defense against viral infection

Researchers decipher how an enzyme modifies the genetic material in the cell nucleus

Large Hadron Collider experiment zeroes in on magnetic monopoles

Scientists discover higher levels of CO₂ increase survival of viruses in the air and transmission risk

Relevant PhysicsForums posts

The Cass Report (UK)

Major Evolution in Action

If theres a 15% probability each month of getting a woman pregnant...

Can four legged animals drink from beneath their feet?

Mold in Plastic Water Bottles? What does it eat?

Dolphins don't breathe through their esophagus

Scientists Map Genetic Regulatory Elements for the Heart

Next gen sequencing technology pinpoint 'on-off switches' in genomes

Exploring the 'last frontier' of our genome

Tracking genes' remote controls

'Moonlighting' molecules discovered

Some Genetic Research is Best Done Close to the Evolutionary Home

Automated machine learning robot unlocks new potential for genetics research

Scientists replace fishmeal in aquaculture with microbial protein derived from soybean processing wastewater

Scientists regenerate neural pathways in mice with cells from rats

Artificial intelligence helps scientists engineer plants to fight climate change

Enhanced CRISPR method enables stable insertion of large genes into the DNA of higher plants

Laser technology offers breakthrough in detecting illegal ivory

Medical Xpress

Tech Xplore

Science X

Computers 'taught' to ID regulating gene sequences

High-precision blood glucose level prediction achieved by few-molecule reservoir computing

Enhancing memory technology: Multiferroic nanodots for low-power magnetic storage

Researchers advance detection of gravitational waves to study collisions of neutron stars and black holes

Automated machine learning robot unlocks new potential for genetics research

AI deciphers new gene regulatory code in plants and makes accurate predictions for newly sequenced genomes

Unveiling a new quantum frontier: Frequency-domain entanglement

Study details a common bacterial defense against viral infection

Researchers decipher how an enzyme modifies the genetic material in the cell nucleus

Large Hadron Collider experiment zeroes in on magnetic monopoles

Scientists discover higher levels of CO₂ increase survival of viruses in the air and transmission risk

Relevant PhysicsForums posts

Related Stories

Scientists Map Genetic Regulatory Elements for the Heart

Next gen sequencing technology pinpoint 'on-off switches' in genomes

Exploring the 'last frontier' of our genome

Tracking genes' remote controls

'Moonlighting' molecules discovered

Some Genetic Research is Best Done Close to the Evolutionary Home

Recommended for you

Automated machine learning robot unlocks new potential for genetics research

Scientists replace fishmeal in aquaculture with microbial protein derived from soybean processing wastewater

Scientists regenerate neural pathways in mice with cells from rats

Artificial intelligence helps scientists engineer plants to fight climate change

Enhanced CRISPR method enables stable insertion of large genes into the DNA of higher plants

Laser technology offers breakthrough in detecting illegal ivory

Newsletter sign up

Donate and enjoy an ad-free experience