Featured Research

from universities, journals, and other organizations

Computers 'taught' to ID regulating gene sequences

Date:
November 5, 2012
Source:
Johns Hopkins Medicine
Summary:
Researchers have succeeded in teaching computers how to identify commonalities in DNA sequences known to regulate gene activity, and to then use those commonalities to predict other regulatory regions throughout the genome. The tool is expected to help scientists better understand disease risk and cell development.

The glowing areas in this zebrafish embryo show the activity of one of the brain enhancer sequences identified. The enhancer is directing the activity of a gene in the lower areas of the central nervous system and in the lens of the eye.
Credit: G. Burzynski

Johns Hopkins researchers have succeeded in teaching computers how to identify commonalities in DNA sequences known to regulate gene activity, and to then use those commonalities to predict other regulatory regions throughout the genome. The tool is expected to help scientists better understand disease risk and cell development.

Related Articles


The work was reported in two recent papers in Genome Research, published online on July 3 and Sept. 27.

"Our goal is to understand how regulatory information is encrypted and to learn which sequence variations contribute to medical risks," says Andrew McCallion, Ph.D., associate professor of molecular and comparative pathobiology in the McKusick-Nathans Institute of Genetic Medicine at Hopkins. "We give data to a computer and 'teach it' to distinguish between data that has no biological value versus data that has this or that biological value. It then establishes a set of rules, which allows it to look at new sets of data and apply what it learned. We're basically sending our computers to school."

These state-of-the-art "machine learning" techniques were developed by Michael Beer, Ph.D., assistant professor of biomedical engineering at the Johns Hopkins School of Medicine, and by Ivan Ovcharenko, Ph.D., at the National Center for Biotechnology Information. The researchers began both studies by creating "training sets" for their computers to "learn" from. These training sets were lists of DNA sequences taken from regions of the genome, called enhancers, that are known to increase the activity of particular genes in particular cells.

For the first of their studies, McCallion's team created a training set of enhancer sequences specific to a particular region of the brain by compiling a list of 211 published sequences that had been shown, by various studies in mice and zebrafish, to be active in the development or function of that part of the brain.

For a second study, the team generated a training set through experiments of their own. They began with a purified population of mouse melanocytes, which are the skin cells that produce the pigment melanin that gives color to skin and absorbs harmful UV rays from the sun. The researchers used a technique called ChIP-seq (pronounced "chip seek") to collect and sequence all of the pieces of DNA that were bound in those cells by special enhancer-binding proteins, generating a list of about 2,500 presumed melanocyte enhancer sequences.

Once the researchers had these two training sets for their computers, one specific to the brain and another to melanocytes, the computers were able to distinguish the features of the training sequences from the features of all other sequences in the genome, and create rules that defined one set from the other. Applying those rules to the whole genome, the computers were able to discover thousands of probable brain or melanocyte enhancer sequences that fit the features of the training sets.

In the brain study, the computers identified 40,000 probable brain enhancer sequences; for melanocytes, 7,500. Randomly testing a subset of each batch of sequences, the scientists found that more than 85 percent of the predicted enhancer sequences enhanced gene activity in the brain or in melanocytes, as expected, verifying the predictive power of their approach.

The researchers say that, in addition to identifying specific DNA sequences that control the genetic activity of a particular organ or cell type, these studies contribute to our understanding of enhancers in general and have validated an experimental approach that can be applied to many other biological questions as well.

Authors on the brain paper include Grzegorz Burzynski, Xylena Reed, Zachary Stine, Takeshi Matsui and Andrew McCallion from The Johns Hopkins University, and Leila Taher and Ivan Ovcharenko from the National Center for Biotechnology Information.

Authors on the melanocyte paper include David Gorkin, Dongwon Lee, Xylena Reed, Christopher Fletez-Brant, Seneca Bessling, Michael Beer and Andrew McCallion from The Johns Hopkins University, and Stacie Loftus and William Pavan from the National Human Genome Research Institute.

This work was supported by grants from the National Institute of Neurological Disorders and Stroke (NS062972), the National Human Genome Research Institute's Intramural Research Program, the National Library of Medicine, the National Institute of General Medical Sciences (GM07814, GM071648), the National Science Foundation and the Searle Scholars Program.


Story Source:

The above story is based on materials provided by Johns Hopkins Medicine. Note: Materials may be edited for content and length.


Journal References:

  1. G. M. Burzynski, X. Reed, L. Taher, Z. E. Stine, T. Matsui, I. Ovcharenko, A. S. McCallion. Systematic elucidation and in vivo validation of sequences enriched in hindbrain transcriptional control. Genome Research, 2012; 22 (11): 2278 DOI: 10.1101/gr.139717.112
  2. D. U. Gorkin, D. Lee, X. Reed, C. Fletez-Brant, S. L. Bessling, S. K. Loftus, M. A. Beer, W. J. Pavan, A. S. McCallion. Integration of ChIP-seq and machine learning reveals enhancers and a predictive regulatory sequence vocabulary in melanocytes. Genome Research, 2012; 22 (11): 2290 DOI: 10.1101/gr.139360.112

Cite This Page:

Johns Hopkins Medicine. "Computers 'taught' to ID regulating gene sequences." ScienceDaily. ScienceDaily, 5 November 2012. <www.sciencedaily.com/releases/2012/11/121105140106.htm>.
Johns Hopkins Medicine. (2012, November 5). Computers 'taught' to ID regulating gene sequences. ScienceDaily. Retrieved October 30, 2014 from www.sciencedaily.com/releases/2012/11/121105140106.htm
Johns Hopkins Medicine. "Computers 'taught' to ID regulating gene sequences." ScienceDaily. www.sciencedaily.com/releases/2012/11/121105140106.htm (accessed October 30, 2014).

Share This



More Computers & Math News

Thursday, October 30, 2014

Featured Research

from universities, journals, and other organizations


Featured Videos

from AP, Reuters, AFP, and other news services

Mind-Controlled Prosthetic Arm Restores Amputee Dexterity

Mind-Controlled Prosthetic Arm Restores Amputee Dexterity

Reuters - Innovations Video Online (Oct. 29, 2014) A Swedish amputee who became the first person to ever receive a brain controlled prosthetic arm is able to manipulate and handle delicate objects with an unprecedented level of dexterity. The device is connected directly to his bone, nerves and muscles, giving him the ability to control it with his thoughts. Matthew Stock reports. Video provided by Reuters
Powered by NewsLook.com
Robots Get Funky on the Dance Floor

Robots Get Funky on the Dance Floor

AP (Oct. 29, 2014) Dancing, spinning and fighting robots are showing off their agility at "Robocomp" in Krakow. (Oct. 29) Video provided by AP
Powered by NewsLook.com
IBM Taps Into Twitter's Data With New Partnership

IBM Taps Into Twitter's Data With New Partnership

Newsy (Oct. 29, 2014) The new partnership will allow IBM to access Twitter’s data and analytics to help IBM clients better understand their consumers. Video provided by Newsy
Powered by NewsLook.com
Google To Use Nanoparticles, Wearables To Detect Disease

Google To Use Nanoparticles, Wearables To Detect Disease

Newsy (Oct. 29, 2014) Google X wants to improve modern medicine with nanoparticles and a wearable device. It's all an attempt to tackle disease detection and prevention. Video provided by Newsy
Powered by NewsLook.com

Search ScienceDaily

Number of stories in archives: 140,361

Find with keyword(s):
Enter a keyword or phrase to search ScienceDaily for related topics and research stories.

Save/Print:
Share:

Breaking News:

Strange & Offbeat Stories


Space & Time

Matter & Energy

Computers & Math

In Other News

... from NewsDaily.com

Science News

Health News

    Environment News

    Technology News



    Save/Print:
    Share:

    Free Subscriptions


    Get the latest science news with ScienceDaily's free email newsletters, updated daily and weekly. Or view hourly updated newsfeeds in your RSS reader:

    Get Social & Mobile


    Keep up to date with the latest news from ScienceDaily via social networks and mobile apps:

    Have Feedback?


    Tell us what you think of ScienceDaily -- we welcome both positive and negative comments. Have any problems using the site? Questions?
    Mobile: iPhone Android Web
    Follow: Facebook Twitter Google+
    Subscribe: RSS Feeds Email Newsletters
    Latest Headlines Health & Medicine Mind & Brain Space & Time Matter & Energy Computers & Math Plants & Animals Earth & Climate Fossils & Ruins